TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published 21 days ago • 18
VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct Paper • 2606.23543 • Published Jun 22 • 6
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents Paper • 2601.18467 • Published Jan 26 • 1
RubricBench: Aligning Model-Generated Rubrics with Human Standards Paper • 2603.01562 • Published Mar 2 • 64