Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Abstract
A shared trajectory model unifies goal and dynamics prediction with action generation for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.
Community
We introduce Dynin-Robotics, an omnimodal masked-diffusion vision-language-action foundation model that unifies policy generation, world modeling, goal-state prediction, and task understanding within a single architecture. By leveraging shared trajectory modeling and iterative bidirectional refinement, Dynin-Robotics integrates visual predictions into action generation and selection, enabling goal-guided control, joint action–future-state denoising, and compositional test-time scaling. As demonstrated in our experiments, Dynin-Robotics achieves competitive performance across simulation benchmarks and real-world manipulation tasks, supporting discrete diffusion as a practical paradigm for unified robotic understanding, prediction, and control.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models (2026)
- TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks (2026)
- SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation (2026)
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning (2026)
- GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation (2026)
- ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training (2026)
- WAM-OPD: On-Policy Distillation for World Action Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.13053 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper