Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Paper • 2607.10169 • Published 19 days ago • 14
view article Article LFM2.5-Encoders for Fast Long-Context Inference on CPU LiquidAI • about 22 hours ago • 45
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness Paper • 2607.08046 • Published 21 days ago • 14
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Paper • 2605.26895 • Published May 26 • 23
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Paper • 2605.25604 • Published May 25 • 139
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation Paper • 2605.11739 • Published May 13 • 61
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR Paper • 2605.15726 • Published May 15 • 36
Efficient Reasoning via Thought-Training and Thought-Free Inference Paper • 2511.03408 • Published Nov 5, 2025 • 1
Why Fine-Tuning Encourages Hallucinations and How to Fix It Paper • 2604.15574 • Published Apr 16 • 26
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published 22 days ago • 24
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Paper • 2607.06987 • Published 22 days ago • 9
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference Paper • 2602.21548 • Published Feb 25 • 56