Research ·
BiMoGen Bidirectional Motion-Text Generation via Masked Discrete Diffusion
AI brief
AI-writtenWhy it mattersProvides new bidirectional generation tech path for motion synthesis and avatar animation.
BiMoGen’s bidirectional framework enables two-way generation between motion and text, outperforming traditional autoregressive approaches in cross-modal consistency.
What happened
The research team introduced BiMoGen, a unified bidirectional motion-text generation framework built on a masked discrete diffusion architecture for bidirectional motion modeling. To ensure stable training, the team designed a decoupled unimodal and cross-modal training pipeline: first, they established cross-modal correspondences on paired motion-text sequences via masked pretraining, then used supervised fine-tuning to adapt the model to bidirectional generation tasks. They also incorporated a generation-aware self-correction mechanism that feeds the model’s own predictions into training, correcting unreliable tokens early in the sampling process to mitigate inference error accumulation common in masked diffusion models.
Key facts
- Framework Name
- BiMoGen
- Core Architecture
- Unified masked discrete diffusion bidirectional motion-text generation framework
- Core Training Design
- Decoupled unimodal/cross-modal training, generation-aware self-correction
- Test Benchmarks
- HumanML3D, KIT-ML
- Project Page
- https://wengwanjiang.github.io/BiMoGen-Page
Background
Text-to-motion generation and motion-to-text annotation are two core foundational tasks in human motion modeling, and both rely on the same set of motion-text correspondences. Most existing unified solutions use autoregressive modeling, which is limited by fixed generation order, struggles to capture the bidirectional dependencies between language and motion, and propagates early prediction errors along the fixed context window, reducing temporal coherence and cross-modal consistency.
Why it matters
For the industry, this framework breaks the sequential limitations of autoregressive modeling for bidirectional motion-text conversion, providing a better cross-modal technical foundation for virtual avatars and motion generation applications. For developers, the framework can simultaneously support motion-driven virtual avatars and automated motion description generation, cutting the maintenance cost of running multiple separate models. For end users, future virtual avatar motion generation will feel more natural and align far better with text instructions.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.