DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer   Normalization Mamba-2

Fan Zhang; Siyuan Zhao; Naye Ji; Zhaohan Wang; Jingmei Wu; Fuxing Gao,; Zhenqing Ye; Leyao Yan; Lanxin Dai; Weidong Geng; Xin Lyu; Bozuo Zhao,; Dingguo Yu; Hui Du; Bin Hu

arXiv:2411.16729·cs.SD·November 28, 2024

DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2

Fan Zhang, Siyuan Zhao, Naye Ji, Zhaohan Wang, Jingmei Wu, Fuxing Gao,, Zhenqing Ye, Leyao Yan, Lanxin Dai, Weidong Geng, Xin Lyu, Bozuo Zhao,, Dingguo Yu, Hui Du, Bin Hu

PDF

Open Access

TL;DR

DiM-Gestor is a novel speech-driven gesture generation model that uses adaptive layer normalization and diffusion techniques to improve efficiency and diversity, especially for Chinese speech-gesture datasets.

Contribution

The paper introduces DiM-Gestor, an end-to-end gesture generation model with adaptive layer normalization, reducing memory usage and increasing inference speed compared to traditional transformer models.

Findings

01

Reduces memory usage by approximately 2.4 times.

02

Increases inference speed by 2 to 4 times.

03

Achieves competitive gesture generation quality on Chinese datasets.

Abstract

Speech-driven gesture generation using transformer-based generative models represents a rapidly advancing area within virtual human creation. However, existing models face significant challenges due to their quadratic time and space complexities, limiting scalability and efficiency. To address these limitations, we introduce DiM-Gestor, an innovative end-to-end generative model leveraging the Mamba-2 architecture. DiM-Gestor features a dual-component framework: (1) a fuzzy feature extractor and (2) a speech-to-gesture mapping module, both built on the Mamba-2. The fuzzy feature extractor, integrated with a Chinese Pre-trained Model and Mamba-2, autonomously extracts implicit, continuous speech features. These features are synthesized into a unified latent representation and then processed by the speech-to-gesture mapping module. This module employs an Adaptive Layer Normalization…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and dialogue systems · Hand Gesture Recognition Systems · Natural Language Processing Techniques

MethodsDiffusion · Layer Normalization