SenseNova-U1 Collection SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture • 12 items • Updated 20 days ago • 76
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 3 days ago • 23
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing Paper • 2608.24777 • Published 10 days ago • 16
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Paper • 2608.08676 • Published 26 days ago • 15
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 29 days ago • 62
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published about 1 month ago • 25
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published Aug 3 • 82
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning Paper • 2606.27771 • Published Jun 26 • 5
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 180
V-JEPA 2 Collection A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of https://ai.meta.com/blog/v-jepa-yann • 8 items • Updated Jun 13, 2025 • 229
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 84
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 77
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Paper • 2604.24763 • Published Apr 27 • 71
Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation Paper • 2604.10030 • Published Apr 11 • 15