open-source-metrics/reinforcement-learning-checkpoint-downloads Viewer • Updated Oct 6, 2022 • 367 • 65 • 4
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 6 days ago • 32
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 6 days ago • 131
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 10 days ago • 150
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 7 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 7 days ago • 42
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 7 days ago • 56
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 8 days ago • 37
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Models Reinforcement Learning • Updated Feb 1 • 76 • 12
Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand Paper • 2609.17172 • Published 9 days ago • 3
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents Paper • 2609.18779 • Published 8 days ago • 14
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 8 days ago • 79
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 8 days ago • 99
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 8 days ago • 62