Towards Open Domain Text-Driven Synthesis of Multi-Person Motions

Mengyi Shan; Lu Dong; Yutao Han; Yuan Yao; Tao Liu; Ifeoma Nwogu,; Guo-Jun Qi; and Mitch Hill

arXiv:2405.18483·cs.CV·July 16, 2024

Towards Open Domain Text-Driven Synthesis of Multi-Person Motions

Mengyi Shan, Lu Dong, Yutao Han, Yuan Yao, Tao Liu, Ifeoma Nwogu,, Guo-Jun Qi, and Mitch Hill

PDF

Open Access

TL;DR

This paper introduces a transformer-based diffusion model capable of generating diverse, high-fidelity multi-person motions from textual descriptions, overcoming dataset limitations and enabling complex group motion synthesis.

Contribution

It presents the first method to generate multi-subject motion sequences from text, utilizing curated datasets and a novel diffusion framework for multi-person motion synthesis.

Findings

01

Successful generation of multi-person static poses.

02

Effective synthesis of multi-person motion sequences.

03

High diversity and fidelity in generated motions.

Abstract

This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Motion and Animation · Human Pose and Action Recognition · Video Analysis and Summarization

MethodsDiffusion