Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 2 days ago • 135
Cliff: Learning Process Rewards from the First Mistake Paper • 2609.02817 • Published 3 days ago • 15
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published 8 days ago • 28
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Paper • 2608.14441 • Published 22 days ago • 29
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 29 days ago • 111
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 24 days ago • 114
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 186
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks Paper • 2606.05445 • Published Jun 3 • 8
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 45
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts Paper • 2509.04500 • Published Sep 2, 2025 • 5
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents Paper • 2412.13549 • Published Dec 18, 2024
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning Paper • 2406.11721 • Published Jun 17, 2024
ShortageSim: Simulating Drug Shortages under Information Asymmetry Paper • 2509.01813 • Published Sep 1, 2025
xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning Paper • 2510.08439 • Published Oct 9, 2025 • 1
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Paper • 2511.02734 • Published Nov 4, 2025 • 23
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering Paper • 2511.13998 • Published Nov 17, 2025 • 3