ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation
Kim Youwang, Lee Hyoseok, Subin Park, Gerard Pons-Moll, Tae-Hyun Oh

TL;DR
ELITE is a novel system that synthesizes high-fidelity, animatable head avatars from monocular videos efficiently by combining learned initialization with test-time generative adaptation, outperforming prior methods in quality and speed.
Contribution
We propose a new Gaussian avatar synthesis framework that integrates a fast initialization model with a test-time diffusion-based refinement, enabling real-time, high-quality avatar generation from monocular videos.
Findings
Achieves 60x faster synthesis than previous 2D generative methods.
Produces superior visual quality and handles challenging expressions.
Demonstrates strong in-the-wild generalization capabilities.
Abstract
We introduce ELITE, an Efficient Gaussian head avatar synthesis from a monocular video via Learned Initialization and TEst-time generative adaptation. Prior works rely either on a 3D data prior or a 2D generative prior to compensate for missing visual cues in monocular videos. However, 3D data prior methods often struggle to generalize in-the-wild, while 2D generative prior methods are computationally heavy and prone to identity hallucination. We identify a complementary synergy between these two priors and design an efficient system that achieves high-fidelity animatable avatar synthesis with strong in-the-wild generalization. Specifically, we introduce a feed-forward Mesh2Gaussian Prior Model (MGPM) that enables fast initialization of a Gaussian avatar. To further bridge the domain gap at test time, we design a test-time generative adaptation stage, leveraging both real and synthetic…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsGenerative Adversarial Networks and Image Synthesis · Face recognition and analysis · Advanced Image Processing Techniques
