StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 7 days ago • 9
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Paper • 2607.05155 • Published Jul 6 • 18
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark Paper • 2607.00724 • Published Jul 1
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 22
MARBLE: Music Audio Representation Benchmark for Universal Evaluation Paper • 2306.10548 • Published Jun 18, 2023
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Paper • 2605.04077 • Published Apr 14 • 7
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation Paper • 2604.02368 • Published Mar 27 • 12
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL Paper • 2602.07078 • Published Feb 6 • 1
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 7 days ago • 9
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 22
OProver: A Unified Framework for Agentic Formal Theorem Proving Paper • 2605.17283 • Published May 17 • 31
InCoder-32B: Code Foundation Model for Industrial Scenarios Paper • 2603.16790 • Published Mar 17 • 312
Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining Paper • 2603.11103 • Published Mar 11 • 9
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization Paper • 2602.22675 • Published Feb 26 • 23