Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 4 days ago • 23
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 9 days ago • 50
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 311
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 235
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories Paper • 2608.08557 • Published 27 days ago • 2
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Paper • 2607.10428 • Published Jul 24 • 1
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Paper • 2606.25325 • Published Jun 24
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy Paper • 2606.27652 • Published Jun 26
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams Paper • 2605.07299 • Published May 8
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories Paper • 2606.11520 • Published Jun 9
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 77
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning Paper • 2605.16371 • Published May 10
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Paper • 2604.15093 • Published Apr 16 • 30