Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published 26 days ago • 148
AsyncOPD: How Stale Can On-Policy Distillation Be? Paper • 2606.24143 • Published about 1 month ago • 31
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published 27 days ago • 111
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 141
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 124
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance Paper • 2604.12627 • Published Apr 14 • 102
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation Paper • 2604.09497 • Published Apr 10 • 29
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings Paper • 2604.04323 • Published Apr 6 • 41