PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published 29 days ago • 95
Parametric Social Identity Injection and Diversification in Public Opinion Simulation Paper • 2603.16142 • Published Jun 1 • 1
Reinforcement Learning from Rich Feedback with Distributional DAgger Paper • 2606.05152 • Published Jun 3 • 3
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents Paper • 2606.05761 • Published Jun 4 • 19
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 44
Ψ-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues Paper • 2606.02754 • Published Jun 1 • 13
GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization Paper • 2605.31464 • Published May 29 • 2