Meridian / docs /inference.md
yycc's picture
Meridian
9f57754
|
Raw
History Blame Contribute Delete
15.4 kB

Inference: camera and time

← Meridian · Installation · Method · Studio demo

Run the examples from the release directory after completing installation. The CLI selects one GPU through CUDA_VISIBLE_DEVICES; it does not split a take across cards.

CUDA_VISIBLE_DEVICES=0 python inference/sample.py \
  --video examples/media/sp_bouldering_hang.mp4 \
  --yaw 15 --sweep --out out/orbit

The default is a 73-frame, 24 fps take using the distilled student: --steps 4 --flow-shift 3. Use a fresh --out directory for each take; filenames are fixed rather than automatically versioned.

Prepare the input

Use a single continuous shot. The CLI does not normalize frame rate or detect cuts: it reads decoded frames by index and always writes at 24 fps. A 30 fps or variable-frame-rate input can therefore change pace and lose audio alignment unless you normalize it first.

# Preserve playback duration while exporting a constant 24 fps input.
ffmpeg -i clip.mp4 -vf "setpts=PTS-STARTPTS,fps=24" \
  -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac clip_24fps.mp4

# Inspect the actual decoded frame count, not only the container's FPS label.
ffprobe -v error -select_streams v:0 -count_frames \
  -show_entries stream=width,height,r_frame_rate,nb_read_frames \
  -of default=noprint_wrappers=1 clip_24fps.mp4

For a normal take, the input must contain at least start + frames decoded frames. The two included sample clips each contain exactly 73 frames at 24 fps and no audio. A longer take needs a longer input or an explicit hold; simply increasing --frames on those samples will not extend the action.

Camera recipes

The following commands all work with the included 73-frame sample as an input, subject to the hardware and model setup. Motion quality still depends on reconstruction and viewpoint coverage.

Orbit with a gentle start and stop

python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --yaw 15 --sweep --ease --out out/eased_orbit

Positive yaw moves the camera left around the pivot. Without --sweep, the offset is applied throughout the clip instead of ramping from the original view. It remains an offset from each source camera, not necessarily a camera fixed in world space.

Push in without changing the lens

python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --dolly 0.8 --zoom 1 --sweep --ease --out out/push_in

--dolly alone performs a dolly zoom: it changes both the camera radius and focal length to approximately preserve the pivot plane's size. Add --zoom 1 for a fixed-lens push-in, where the subject grows in frame. --zoom 1.5 without translation is an optical zoom; these pure-zoom takes can be ignored by the model. Prefer moves with parallax.

Slide or crane while keeping the subject framed

# Move right by 0.15 pivot-depth units and aim back toward the pivot.
python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --truck 0.15 --aim --sweep --ease --out out/slide

# Raise the camera by 0.15 pivot-depth units.
python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --boom 0.15 --aim --sweep --ease --out out/crane

Choose the orbit center

python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --pivot 0.5,0.5 --pivot-lock --yaw 15 --sweep --out out/pivot_orbit

--pivot fx,fy uses fractions of the center-cropped picture, not the full letterbox or original uncropped image. Choose a point on the subject, away from image boundaries, sky, and missing depth. The example selects the crop center; adjust it for your footage. --pivot sets the depth scale; --pivot-lock additionally moves the orbit center to the picked 3D point.

Choose the point in the frame used for pivot depth. Without --freeze, this is the first selected source frame (--start). With --freeze F:N, it is frame F. An athlete-centered point at the held apex can land on the distant audience in the approach frame: do not reuse it unchanged when removing the hold. The depth is a median over a neighborhood extending roughly 5% of the picture in each direction, so check that neighborhood as well as the exact pixel. --pivot-lock is not dynamic subject tracking; inspect the projected reference throughout the shot before treating the requested trajectory as a successful composition.

Timing

All CLI source indices are zero-based absolute frame indices in the supplied file.

Select a passage

# On an input with at least 121 frames, use source frames 48 through 120 inclusive.
python inference/sample.py --video clip_24fps.mp4 \
  --start 48 --frames 73 --yaw 15 --sweep --out out/later_moment

--start 48 is two seconds into a 24 fps input. It is not a seek time in seconds.

Hold a moment while moving the camera

# 24 live frames, then source frame 24 repeated for 49 output frames; no tail.
python inference/sample.py --video examples/media/sp_bouldering_reach.mp4 \
  --yaw 35 --freeze 24:49 --out out/bullet

# Hold one instant for the entire take; the camera still sweeps through 20 degrees.
python inference/sample.py --video examples/media/sp_bouldering_reach.mp4 \
  --yaw 20 --freeze 24:73 --start 24 --out out/held_moment

For --freeze F:N, the output consists of:

  1. Source frames start through F - 1, once each.
  2. Source frame F, repeated N times.
  3. Source frames after F, once each, until the requested output length is reached.

The tail length is frames - N - (F - start). Use a positive N, F >= start, and a non-negative tail. The source must reach frame F + tail. With a hold, fewer distinct source frames can produce a longer output, but the repeated interval contains no new event motion.

By default, the camera ramp progresses only during the held interval. Add --sweep to move during the live lead-in and tail too; --live-speed sets their ramp speed relative to the held interval (default 0.33). It controls the camera ramp, not playback speed.

The CLI can also decode a still image, for example --video still.png --freeze 0:73 --yaw 15. This is a camera move over a held image, not animation of the subject. The studio's video upload workflow does not offer this short-input path.

Slow motion and speed-ups

Set the event's pace first, then author the camera. Export the retimed source at a constant 24 fps and use that export as the input to either the CLI or the studio.

# 0.5x: twice the duration.
ffmpeg -i clip.mp4 -vf "setpts=2*(PTS-STARTPTS),fps=24" -an clip_slow.mp4

# 2x: half the duration.
ffmpeg -i clip.mp4 -vf "setpts=0.5*(PTS-STARTPTS),fps=24" -an clip_fast.mp4

python inference/sample.py --video clip_slow.mp4 \
  --yaw 15 --sweep --out out/slow_orbit

python inference/sample.py --video clip_fast.mp4 \
  --yaw 15 --sweep --out out/fast_orbit

Changing playback metadata alone is not enough for the CLI: timing must be baked into the decoded frame sequence. fps=24 duplicates or drops frames; it does not interpolate new motion. Slow-motion smoothness depends on the source frame rate and any interpolation applied before inference.

Check the retimed file's frame count before rendering. In particular, acceleration shortens the input: the included 73-frame samples become too short for an ordinary 73-frame take at 2x. Use longer footage or explicitly hold a moment. The commands above omit audio; retime a soundtrack separately if needed.

In the studio, one source frame per output frame preserves the export's retimed pace. Stretching keys across the unchanged source is a different workflow: it duplicates or skips reconstructed frames and can trigger a speed warning.

Output lengths and resolution

--frames Duration at 24 fps
73 3.04 s
90 3.75 s
107 4.46 s
124 5.17 s
141 5.88 s
158 6.58 s
175 7.29 s
243 10.13 s

These are the lengths with shipped text and audio-layout assets. Other values are not accepted by the CLI. This is a per-take limit, not a limit on the total duration of the input file.

Output uses an aspect-matched 768-class bucket, usually about 1.03 million pixels; 16:9 maps to 1344 × 768. Both references use the smaller 480-class bucket: 832 × 480 for a 16:9 input.

Output files

File in --out Contents
out.mp4 Generated take, 24 fps, no generated audio.
render.mp4 Geometry reference at conditioning resolution, including grey holes.
source.mp4 Source images after the selected frame mapping and output crop. A hold is visible here too.
grid.mp4 Source, geometry reference, and generated take side by side.
out_audio.mp4 For non-freeze commands: source-window audio muxed onto the take, when the input has audio. A silent input remains silent.
last.png Final generated frame. Reusing it is possible, but does not guarantee cross-take consistency.
cams.npz Source/target camera matrices, intrinsics, pivot metadata, crop, canvas, FPS, and command arguments.

cams.npz stores c2w_src and c2w_dst as camera-to-world matrices; intr_src and intr_dst are in the 512-space geometry grid. The *_px arrays are exported for the output canvas. Translation units are reconstruction-relative, not meters. The archive is diagnostic metadata, not a scene model.

For reproducibility, retain the input export, command, seed, checkpoint revisions, and environment. Do not assume the CLI and studio, or different dependency/backend versions, produce bit-identical results from the same seed.

CLI reference

Run python inference/sample.py --help for the parser's complete help. The tables below group the options by purpose; boolean flags are off unless stated otherwise.

Input and model

Option Default Meaning
--video Required Input video or decodable still image.
--out Required Output directory.
--start 0 First source-frame index.
--frames 73 Supported output length from the table above.
--seed 1234 Random seed.
--ckpt <release>/transformer Finetuned teacher directory.
--lora <release>/lora Student adapter directory.
--no-lora Off Disable the adapter; pair with the teacher sampling settings.
--steps 4 Scheduler grid points, including the terminal point.
--flow-shift 3 Video schedule shift; use 12 for the teacher.
--model-dir MiniMaxAI/MiniMax-H3 Hub repo or local directory containing vae/.
--vggt-repo, --vggt Environment-based VGGT-Omega source checkout and checkpoint; see Installation.
--attn-backend _native_cudnn Diffusers attention backend. Alternatives are hardware- and version-dependent.

To use the teacher:

python inference/sample.py --video examples/media/sp_bouldering_hang.mp4 \
  --yaw 15 --sweep --no-lora --steps 50 --flow-shift 12 --out out/teacher

Camera and motion

Option Default Meaning
--yaw 0 Orbit angle in degrees; positive moves left.
--yaw-from 0 Initial yaw when ramping; the live lead-in holds this value unless also swept.
--truck 0 Sideways shift in pivot-depth units; positive moves right.
--boom 0 Vertical shift in pivot-depth units; positive raises the camera.
--dolly 1 Orbit-radius scale; below one moves closer and, by default, widens the lens.
--zoom 0 Zero means automatic dolly-linked focal scaling; a positive value specifies the final focal multiplier.
--pivot Unset fx,fy in the crop; selects the depth-scale neighborhood.
--pivot-lock Off With --pivot, orbit about the selected 3D point.
--aim Off Reorient toward the pivot after translation.
--pivot-to Unset With --aim, a second fx,fy point toward which the aim transitions.
--sweep Off Ramp from the initial to the final offset over the take.
--ease Off Cosine ease-in/out applied to the ramp; it does not create a ramp by itself.
--bounce Off There-and-back ramp, 0 → 1 → 0; implies a ramp even without --sweep.
--swing Off Sine ramp, 0 → 1 → 0 → −1 → 0; implies a ramp.
--freeze Unset F:N: hold source frame F for N output frames.
--live-speed 0.33 With --freeze --sweep, relative camera-ramp speed outside the hold.

Use one basic ramp shape at a time. Combining --ease, --bounce, and --swing composes their functions in code order; it does not select between independent motion presets.

Advanced and diagnostic controls

Option Default Meaning and caveat
--gauge-only Off Reconstruct, warp, print geometry gauges, then stop before loading H3. No normal output artifacts are written.
--follow Off Replay estimated source cameras over the first selected frame's fixed geometry and RGB. Ignores the authored yaw/translation/lens controls; it does not retain the event's live motion.
--smooth 8 With --follow, Gaussian smoothing sigma in frames for estimated camera poses and intrinsics. 0 disables smoothing.
--cull Off Reject surfaces seen from behind according to estimated depth-map normals. This removes misleading splats; it does not reveal hidden surfaces.
--fast-back 1 Above one, compress the middle half of the camera ramp. Does not improve the reconstruction of an unseen back view.
--canvas Automatic Explicit WxH, with dimensions divisible by 32; advanced override outside the reported default benchmarks.
--full 0 → 1280 Override the square letterbox side. Higher values increase point-cloud sampling and memory, not VGGT's 512-pixel input resolution.

The CLI prints ahead, coverage, and other geometry diagnostics but does not reject a take using the studio's clearance/motion thresholds. Inspect the diagnostics and render.mp4; do not treat a successful process exit as a quality check.

Improving a take

  1. Inspect the geometry reference first. A bent subject or unstable depth in render.mp4 usually needs a better source shot or a smaller move, not more denoising steps.
  2. Use modest viewpoint changes. Large orbits reveal surfaces absent from the source. A plausible completion can still be wrong; roughly 40° is a reported caution point, not a universal threshold.
  3. Check the depth scale. If a small numerical move sends the camera through the scene, pick a pivot on the subject and reduce the translation.
  4. Add parallax to lens changes. Use a small dolly rather than relying on a pure optical zoom.
  5. Check timing before inference. Stuttering from repeated input frames is not a geometry failure; use higher-frame-rate footage or an interpolated export when smooth slow motion matters.
  6. Do not cross cuts. Split the input into continuous shots yourself when using the CLI.

For dependency and memory errors, see Setup problems.