Research ·

BiMoGen Bidirectional Motion-Text Generation via Masked Discrete Diffusion

61Developing1 reportHF Daily Papers
BiMoGen:统一掩码离散扩散的动作文本双向生成
Image: HF Daily Papers

AI brief

AI-written

Why it mattersProvides new bidirectional generation tech path for motion synthesis and avatar animation.

BiMoGen’s bidirectional framework enables two-way generation between motion and text, outperforming traditional autoregressive approaches in cross-modal consistency.

What happened

The research team introduced BiMoGen, a unified bidirectional motion-text generation framework built on a masked discrete diffusion architecture for bidirectional motion modeling. To ensure stable training, the team designed a decoupled unimodal and cross-modal training pipeline: first, they established cross-modal correspondences on paired motion-text sequences via masked pretraining, then used supervised fine-tuning to adapt the model to bidirectional generation tasks. They also incorporated a generation-aware self-correction mechanism that feeds the model’s own predictions into training, correcting unreliable tokens early in the sampling process to mitigate inference error accumulation common in masked diffusion models.

Key facts

Framework Name
BiMoGen
Core Architecture
Unified masked discrete diffusion bidirectional motion-text generation framework
Core Training Design
Decoupled unimodal/cross-modal training, generation-aware self-correction
Test Benchmarks
HumanML3D, KIT-ML
Project Page
https://wengwanjiang.github.io/BiMoGen-Page

Background

Text-to-motion generation and motion-to-text annotation are two core foundational tasks in human motion modeling, and both rely on the same set of motion-text correspondences. Most existing unified solutions use autoregressive modeling, which is limited by fixed generation order, struggles to capture the bidirectional dependencies between language and motion, and propagates early prediction errors along the fixed context window, reducing temporal coherence and cross-modal consistency.

Why it matters

For the industry, this framework breaks the sequential limitations of autoregressive modeling for bidirectional motion-text conversion, providing a better cross-modal technical foundation for virtual avatars and motion generation applications. For developers, the framework can simultaneously support motion-driven virtual avatars and automated motion description generation, cutting the maintenance cost of running multiple separate models. For end users, future virtual avatar motion generation will feel more natural and align far better with text instructions.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. HF Daily Papers ↗BiMoGen Bidirectional Motion-Text Generation via Masked Discrete DiffusionBiMoGen based on unified masked discrete diffusion enables bidirectional motion-text generation.
Back to AI News