WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics Paper • 2603.13391 • Published Mar 11 • 21
On the Design Fundamentals of Pixel Text Representation Learning Paper • 2609.01147 • Published 2 days ago • 17
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation Paper • 2608.29846 • Published 4 days ago • 10
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 1 day ago • 5
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 2 days ago • 13
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 202
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 6 days ago • 100
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 3 days ago • 47
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 20 days ago • 282
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 24 days ago • 342
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 10 days ago • 64
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction Paper • 2605.29341 • Published May 28 • 20
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 10 days ago • 205
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 13 days ago • 63
Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses Paper • 2608.08466 • Published 25 days ago • 13
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published 15 days ago • 121
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 14 days ago • 274
Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling Paper • 2501.11651 • Published Jan 20, 2025 • 2
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127