TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 3 days ago • 143
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published 9 days ago • 13
stabilityai/stable-video-diffusion-img2vid-xt Image-to-Video • 2B • Updated Jul 10, 2024 • 202k • 3.36k
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising Paper • 2607.00407 • Published 21 days ago • 10
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
Rethinking the Role of Efficient Attention in Hybrid Architectures Paper • 2606.15378 • Published Jun 13 • 20
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 113
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 128
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent Paper • 2605.24468 • Published May 23 • 9
Representation over Routing: Diagnosing Temporal Routing Pathologies in Multi-Timescale PPO Paper • 2604.13517 • Published May 30 • 5