3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control
Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, Naoya Chiba, Yuki Uranishi

TL;DR
3DGesPolicy introduces a novel action-based diffusion framework for generating coherent, expressive, and speech-aligned holistic co-speech gestures by modeling motion as unified trajectories and integrating multi-modal signals.
Contribution
It reformulates gesture generation as a continuous trajectory control problem and proposes a multi-modal fusion module for better semantic and expressive alignment.
Findings
Outperforms state-of-the-art methods in naturalness and expressiveness
Ensures spatial and semantic coherence in gesture generation
Effectively models inter-frame motion patterns
Abstract
Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or frame-level regression methods, We introduce 3DGesPolicy, a novel action-based framework that reformulates holistic gesture generation as a continuous trajectory control problem through diffusion policy from robotics. By modeling frame-to-frame variations as unified holistic actions, our method effectively learns inter-frame holistic gesture motion patterns and ensures both spatially and semantically coherent movement trajectories that adhere to realistic motion manifolds. To further bridge the gap in expressive alignment, we propose a Gesture-Audio-Phoneme (GAP) fusion module that can deeply integrate and refine multi-modal signals, ensuring…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsSocial Robot Interaction and HRI · Human Motion and Animation · Face recognition and analysis
