RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems Paper • 2609.09657 • Published 12 days ago • 4
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 25 days ago • 35
CL4SE: A Context Learning Benchmark For Software Engineering Tasks Paper • 2602.23047 • Published Feb 26 • 2
Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs Paper • 2409.10033 • Published Sep 16, 2024 • 1
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning Paper • 2506.03136 • Published Jun 3, 2025 • 25