Unsupervised Cross-Modal Alignment for Multi-Person 3D Pose Estimation

Jogendra Nath Kundu; Ambareesh Revanur; Govind Vitthal Waghmare; Rahul; Mysore Venkatesh; R. Venkatesh Babu

arXiv:2008.01388·cs.CV·August 5, 2020

Unsupervised Cross-Modal Alignment for Multi-Person 3D Pose Estimation

Jogendra Nath Kundu, Ambareesh Revanur, Govind Vitthal Waghmare, Rahul, Mysore Venkatesh, R. Venkatesh Babu

PDF

1 Repo

TL;DR

This paper introduces a fast, deployment-friendly bottom-up framework for multi-person 3D pose estimation that learns a shared latent space through cross-modal alignment, eliminating the need for paired supervision and keypoint grouping.

Contribution

It proposes a novel neural representation for multi-person 3D poses, a cross-modal training paradigm without paired annotations, and a generative pose embedding that improves speed and accuracy.

Findings

01

Achieves state-of-the-art results among bottom-up methods.

02

Generalizes well to in-the-wild images.

03

Offers a superior speed-performance trade-off.

Abstract

We present a deployment friendly, fast bottom-up framework for multi-person 3D human pose estimation. We adopt a novel neural representation of multi-person 3D pose which unifies the position of person instances with their corresponding 3D pose representation. This is realized by learning a generative pose embedding which not only ensures plausible 3D pose predictions, but also eliminates the usual keypoint grouping operation as employed in prior bottom-up approaches. Further, we propose a practical deployment paradigm where paired 2D or 3D pose annotations are unavailable. In the absence of any paired supervision, we leverage a frozen network, as a teacher model, which is trained on an auxiliary task of multi-person 2D pose estimation. We cast the learning as a cross-modal alignment problem and propose training objectives to realize a shared latent space between two diverse modalities.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

revanurambareesh/multiperson
tf

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsDogecoin Customer Service Number +1-833-534-1729