Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos

Matthew Chang; Aditya Prakash; Saurabh Gupta

arXiv:2305.16301·cs.CV·May 26, 2023·1 cites

Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos

Matthew Chang, Aditya Prakash, Saurabh Gupta

PDF

Open Access 1 Video

TL;DR

This paper introduces a method to separate human hand and environment in egocentric videos, improving robotic task performance by using a diffusion model for better scene inpainting and representation.

Contribution

The work presents a novel factorization approach and a diffusion-based inpainting model that enhances scene understanding and downstream robotic applications from egocentric videos.

Findings

01

VIDM improves inpainting quality in egocentric videos

02

Factored representation aids object detection and 3D reconstruction

03

Enhances learning of reward functions and policies from videos

Abstract

The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents a nuisance. However, often hands also provide a valuable signal, e.g. the hand pose may suggest what kind of object is being held. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos· slideslive

Taxonomy

TopicsGenerative Adversarial Networks and Image Synthesis · Advanced Vision and Imaging · Face recognition and analysis

MethodsDiffusion · Inpainting