Realistic Lip Motion Generation Based on 3D Dynamic Viseme and Coarticulation Modeling for Human-Robot Interaction

Sheng Li; Jingcheng Huang; and Min Li

arXiv:2604.01756·cs.RO·April 3, 2026

Realistic Lip Motion Generation Based on 3D Dynamic Viseme and Coarticulation Modeling for Human-Robot Interaction

Sheng Li, Jingcheng Huang, and Min Li

PDF

1 Repo

TL;DR

This paper introduces a novel framework for generating realistic lip motion for humanoid robots using 3D viseme and coarticulation modeling, validated through experiments and available code.

Contribution

It presents a new lip motion generation method based on 3D viseme and coarticulation modeling tailored for humanoid robots, with practical deployment.

Findings

01

High correlation between generated and real lip motions (PCC)

02

Reduced jerk in lip movements (MAJ)

03

Effective retargeting to 14-DOF humanoid lip system

Abstract

Realistic lip synchronization is essential for the natural human-robot non-verbal interaction of humanoid robots. Motivated by this need, this paper presents a lip motion generation framework based on 3D dynamic viseme and coarticulation modeling. By analyzing Chinese pronunciation theory, a 3D dynamic viseme library is constructed based on the ARKit standard, which offers coherent prior trajectories of lips. To resolve motion conflicts within continuous speech streams, a coarticulation mechanism is developed by incorporating initial-final (Shengmu-Yunmu) decoupling and energy modulation. After developing a strategy to retarget high-dimensional spatial lip motion to a 14-DOF lip actuation system of a humanoid head platform, the efficiency and accuracy of the proposed architecture is experimentally validated and demonstrated with quantitative ablation experiments using the metrics of the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yuesheng21/Phoneme-to-Lip-14DOF
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.