Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

Jingi Kim; Wonjun Kim

arXiv:2604.25466·cs.CV·April 29, 2026

Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

Jingi Kim, Wonjun Kim

PDF

TL;DR

This paper introduces a novel approach for human Gaussian splatting that leverages multi-view semantic consistency to improve 3D Gaussian localization and rendering quality from sparse views.

Contribution

It proposes a new method using cross-view attention to unproject and recalibrate latent embeddings, addressing feature inconsistency across views.

Findings

01

Improves 3D Gaussian localization accuracy.

02

Enhances human rendering quality from sparse views.

03

Outperforms existing methods on benchmark datasets.

Abstract

Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing methods rely on explicit geometric constraints or predefined structural representations to accurately position 3D Gaussians. Although these approaches have shown the remarkable progress in this field, they still suffer from inconsistent feature representations across multi-view inputs due to complex articulations of the human body and limited overlaps between different views. To address this problem, we propose a novel method to accurately localize 3D Gaussians and ultimately improve the quality of human rendering. The key idea is to unproject latent embeddings encoded from each viewpoint into a shared 3D space through predicted depth maps and recalibrate them belonging to the same body part based on cross-view attention. This helps the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.