MotionHiFlow: Text-to-motion via hierarchical flow matching

Heng Li; Xiaotong Lin; Ling-An Zeng; Yulei Kang; Shuai Li; Jian-Fang Hu

arXiv:2604.23264·cs.CV·April 28, 2026

MotionHiFlow: Text-to-motion via hierarchical flow matching

Heng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang, Shuai Li, Jian-Fang Hu

PDF

1 Repo

TL;DR

MotionHiFlow introduces a hierarchical flow matching framework for text-to-motion generation, capturing multi-scale motion semantics and details to produce more coherent and semantically aligned 3D human motions.

Contribution

The paper proposes a novel hierarchical flow matching approach that models motions across multiple temporal scales, improving semantic alignment and motion detail in text-to-motion generation.

Findings

01

Achieves state-of-the-art results on HumanML3D and KIT-ML benchmarks.

02

Demonstrates the effectiveness of hierarchical design through ablation studies.

03

Integrates a Text-Motion Diffusion Transformer and topology-aware Motion VAE for better structural modeling.

Abstract

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which limits both semantic alignment and temporal coherence. Inspired by the fact that complex motions are conceptualized hierarchically rather than at a single temporal scale in the human cognitive system, we propose \textit{MotionHiFlow}, a hierarchical flow matching framework to generate motion progressively by constructing flow path from low to high temporal scales. The flows at lower scales capture high-level semantics and coarse motion structures, while flows at higher scales refine temporal details. To link the flows across scales, we introduce a novel cross-scale transition process,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ai-lh/MotionHiFlow
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.