Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 7 days ago • 28
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 7 days ago • 142
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 14 days ago • 64
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 17 days ago • 63
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 18 days ago • 111
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published 13 days ago • 70
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 14 days ago • 205
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 32
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning Paper • 2012.13255 • Published Dec 22, 2020 • 6
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers Paper • 2106.04647 • Published Jun 8, 2021 • 2
Training language models to follow instructions with human feedback Paper • 2203.02155 • Published Mar 4, 2022 • 26
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Paper • 2205.14135 • Published May 27, 2022 • 16