EMAGE Model Card

EMAGE is the CVPR 2024 work Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling. It generates a holistic SMPL-X sequence—face, body, hands, and global translation—from speech audio. Motius wraps the official PantoMatrix audio checkpoint and its five released compositional decoders.

Release Snapshot

Item Value
Task Speech-to-Gesture
Dataset BEAT2 English speaker 2
Input 16 kHz mono waveform
Native motion SMPL-X 55-joint axis-angle pose + expression + translation
Motion rate 30 FPS
Window 64 frames with four autoregressive seed frames
Checkpoint ZeyuLing/Motius-EMAGE-BEAT2
Upstream revision H-Liu1997/emage_audio@b04466b83bdebd656844a2c416797d3131e34ede
Runtime revision PantoMatrix/PantoMatrix@c7356f35f8e39e469e510ccd1bf37e44adf8ec0e
License Apache-2.0

Setup

The released PyTorch .bin files require the compatible Transformers loader used by the official model:

bash tools/setup_emage_env.sh

The PantoMatrix inference runtime is optional and remains attributable to its official source:

git clone https://github.com/PantoMatrix/PantoMatrix ref_repo/PantoMatrix
git -C ref_repo/PantoMatrix checkout c7356f35f8e39e469e510ccd1bf37e44adf8ec0e
export MOTIUS_EMAGE_ROOT="$PWD/ref_repo/PantoMatrix"

Usage

from motius.pipelines.emage import EMAGEPipeline

pipeline = EMAGEPipeline.from_pretrained(
    "ZeyuLing/Motius-EMAGE-BEAT2",
    bundle_kwargs={
        "runtime_root": "ref_repo/PantoMatrix",
        "device": "cuda",
    },
)
motion = pipeline.infer_s2g(["speech.wav"])[0]
pipeline.save_npz(
    motion,
    "outputs/inference/emage/speech_to_gesture/speech.npz",
)

The output NPZ retains all 55 SMPL-X joints, 100 expression coefficients, global translation, neutral betas, and the 30 FPS rate. Motius replaces one upstream CUDA-only tensor allocation with a device-preserving equivalent; the network, checkpoint, chunking, autoregressive seed, and decoder semantics are unchanged.

Evaluation Results

Motius uses the same fixed 15-clip BEAT2 English speaker-2 population and official feature/onset backends as the Language of Motion row. Predictions are trimmed to the paired reference frame count only after the official 64-frame autoregressive inference path.

Evaluator Clips FGD ↓ BC ↑ Diversity uTMR FID ↓ uTMR Paired Dist. ↓
BEAT2 Official 15 6.199 7.564 12.476 0.0523 17.1924

The released audio checkpoint measures FGD = 6.199, BC = 7.564, and Diversity = 12.476 on the paper display scale. The paper reports 5.512 / 7.724 / 13.06; the relative differences are +12.46% / -2.07% / -4.47% respectively. This is official-checkpoint verification, not an exact paper reproduction: the maintained public checkpoint is audio-only, while the paper model also consumed text. The uTMR columns use only canonical 30 FPS SMPL-22 joint positions; metric FK sets SMPL-X translation to zero. FID is L2-normalized, while paired distance uses native uTMR embeddings.

Preview

Open the audio-synchronized 15-case Three.js comparison. It compares BEAT2 GT, Language of Motion, EMAGE, and DiffuseStyleGesture+ with shared camera and timeline controls and native NPZ downloads.

Citation

@inproceedings{liu2024emage,
  title={EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling},
  author={Liu, Haiyang and Zhu, Zihao and Becherini, Giorgio and Peng, Yichen and Su, Mingyang and Zhou, You and Zhe, Xuefei and Iwamoto, Naoya and Zheng, Bo and Black, Michael J.},
  booktitle={CVPR},
  year={2024}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZeyuLing/Motius-EMAGE-BEAT2

Finetuned
(1)
this model

Paper for ZeyuLing/Motius-EMAGE-BEAT2