HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes

Zan Wang; Yixin Chen; Tengyu Liu; Yixin Zhu; Wei Liang; Siyuan Huang

arXiv:2210.09729·cs.CV·October 19, 2022·23 cites

HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes

Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, Siyuan Huang

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces HUMANISE, a large-scale, semantic-rich synthetic dataset for language-conditioned human motion generation in 3D scenes, and proposes a model to generate diverse, scene-aware motions based on natural language descriptions.

Contribution

The paper creates a new large-scale dataset with semantic annotations and develops a novel model for language-conditioned human motion generation in 3D scenes.

Findings

01

The model generates diverse, scene-aware human motions.

02

The dataset enables better semantic understanding of human-scene interactions.

03

Experiments show the model's effectiveness in producing realistic motions.

Abstract

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale and semantic-rich synthetic HSI dataset, denoted as HUMANISE, by aligning the captured human motion sequences with various 3D indoor scenes. We automatically annotate the aligned motions with language descriptions that depict the action and the unique interacting objects in the scene; e.g., sit on the armchair near the desk. HUMANISE thus enables a new generation task, language-conditioned human motion generation in 3D scenes. The proposed task is challenging as it requires joint modeling of the 3D scene, human motion, and natural language. To tackle this task, we present a novel…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Silverster98/HUMANISE
pytorchOfficial

Videos

HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes· slideslive

Taxonomy

TopicsHuman Pose and Action Recognition · Human Motion and Animation · Multimodal Machine Learning Applications