3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Xuanmeng Sha; Liyun Zhang; Tomohiro Mashita; Naoya Chiba; Yuki Uranishi

arXiv:2409.10848·cs.CV·August 13, 2025

3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, Naoya Chiba, Yuki Uranishi

PDF

Open Access

TL;DR

3DFacePolicy introduces an action-based control paradigm for audio-driven 3D facial animation, enabling more natural and expressive movements by predicting vertex action sequences conditioned on audio.

Contribution

It pioneers a new approach by defining vertex trajectories through actions and employing a diffusion policy for improved animation quality.

Findings

01

Outperforms state-of-the-art methods on VOCASET and BIWI datasets.

02

Produces more dynamic, expressive, and smooth facial animations.

03

Effective in generating natural and continuous facial movements.

Abstract

Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsFace recognition and analysis

MethodsDiffusion · Focus