Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Paper • 2512.05774 • Published Dec 5, 2025 • 7
Learning Visual Grounding from Generative Vision and Language Model Paper • 2407.14563 • Published Jul 18, 2024
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Paper • 2607.05390 • Published 23 days ago • 11
A Cookbook of 3D Vision: Data, Learning Paradigms, and Application Paper • 2606.04291 • Published Jun 2 • 5
Causal-JEPA: Learning World Models through Object-Level Latent Interventions Paper • 2602.11389 • Published Feb 11 • 12
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models Paper • 2601.05376 • Published Jan 8 • 1
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals Paper • 2505.19386 • Published May 26, 2025 • 11
Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals Paper • 2601.05848 • Published Jan 9 • 16
Text2Zinc: A Cross-Domain Dataset for Modeling Optimization and Satisfaction Problems in MiniZinc Paper • 2503.10642 • Published Feb 22, 2025 • 2
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models Paper • 2506.11110 • Published Jun 8, 2025 • 4
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals Paper • 2505.19386 • Published May 26, 2025 • 11