Safin-1: Safety from Within through Memory-Native State Evolution Paper • 2609.00092 • Published 4 days ago • 20
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Paper • 2608.12435 • Published 23 days ago • 1
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Paper • 2608.23564 • Published 11 days ago • 14
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement Paper • 2608.20318 • Published 15 days ago • 3
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation Paper • 2603.22193 • Published Apr 2
CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations Paper • 2512.23328 • Published Jan 1
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm Paper • 2509.23946 • Published Sep 28, 2025
Safin-1: Safety from Within through Memory-Native State Evolution Paper • 2609.00092 • Published 4 days ago • 20
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 150
Improving Sampling for Masked Diffusion Models via Information Gain Paper • 2602.18176 • Published Mar 18 • 2
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Paper • 2604.12290 • Published Apr 14 • 16
Improving Sampling for Masked Diffusion Models via Information Gain Paper • 2602.18176 • Published Mar 18 • 2
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Paper • 2604.12290 • Published Apr 14 • 16
Light of Normals: Unified Feature Representation for Universal Photometric Stereo Paper • 2506.18882 • Published Jun 23, 2025 • 89