StereoWorld

Official model weights for StereoWorld: Camera-Guided Stereo Video Generation.

GitHub Paper

Models

Directory Model Description
StereoWorldModel/ Fixed-Baseline Stereo Generates side-by-side stereo video with a configurable but fixed stereo baseline. Use --use_raymap during inference.
StereoWorldFlexModel/ Flexible Stereo Provides independent left/right camera control with converging, horizontal-offset, depth-offset, and height-offset right-camera modes.
StereoWorldInpaintModel/ Fixed-Left View Inpainting Generates a right-view video from an input left-view video while locking the left-video latents during sampling.

Fixed-Baseline Stereo (StereoWorldModel/)

A binocular teacher model for standard stereo video generation with consistent disparity and a fixed baseline.

Flexible Stereo (StereoWorldFlexModel/)

A multi-view world model with independently controlled left and right camera trajectories and four flexible right-camera modes.

Fixed-Left View Inpainting (StereoWorldInpaintModel/)

Generates a right-view video from an input left-view video with camera-guided fixed-left view inpainting.

Download

Download everything

huggingface-cli download Yang-Tian/StereoWorld --local-dir weights

Download only Fixed-Baseline Stereo

huggingface-cli download Yang-Tian/StereoWorld \
  --include "StereoWorldModel/*" \
  --local-dir weights

Download only Flexible Stereo

huggingface-cli download Yang-Tian/StereoWorld \
  --include "StereoWorldFlexModel/*" \
  --local-dir weights

Download only Fixed-Left View Inpainting

huggingface-cli download Yang-Tian/StereoWorld \
  --include "StereoWorldInpaintModel/*" \
  --local-dir weights

Directory Structure

StereoWorld/
β”œβ”€β”€ StereoWorldModel/
β”‚   β”œβ”€β”€ transformer/
β”‚   β”œβ”€β”€ vae/
β”‚   β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ text_encoder/
β”‚   └── scheduler/
β”œβ”€β”€ StereoWorldFlexModel/
β”‚   β”œβ”€β”€ transformer/
β”‚   β”œβ”€β”€ vae/
β”‚   β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ text_encoder/
β”‚   └── scheduler/
└── StereoWorldInpaintModel/
    β”œβ”€β”€ transformer/
    β”œβ”€β”€ vae/
    β”œβ”€β”€ tokenizer/
    β”œβ”€β”€ text_encoder/
    └── scheduler/

Usage

Clone the StereoWorld code repository and install its dependencies before running inference:

git clone https://github.com/SunYangtian/StereoWorld.git
cd StereoWorld
pip install -r requirements.txt

Fixed-Baseline Stereo

python3 inference.py \
  --pipeline_dir weights/StereoWorldModel \
  --use_raymap \
  --eval_json ExpData/demo_custom_eval.json

Flexible Stereo

python3 inference_flex.py \
  --pipeline_dir weights/StereoWorldFlexModel \
  --eval_json ExpData/flex_demo_custom_eval.json

Fixed-Left View Inpainting

python3 inference_view_inpainting.py \
  --pipeline_dir weights/StereoWorldInpaintModel \
  --eval_json ExpData/view_inpaint_eval.json

Outputs for the bundled examples are organized under output_view_inpaint/caseN/.

See the GitHub repository for installation details and further inference options.

Citation

@article{sun2026stereo,
  title={Stereo World Model: Camera-Guided Stereo Video Generation},
  author={Sun Yang-Tian and Huang Zehuan and Niu Yifan and Ma Lin and Cao Yan-Pei and Ma Yuewen and Qi Xiaojuan},
  journal={arXiv preprint arXiv:2603.17375},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for Yang-Tian/StereoWorld