Do Egocentric Video-Language Models Truly Understand Hand-Object   Interactions?

Boshen Xu; Ziheng Wang; Yang Du; Zhinan Song; Sipeng Zheng; Qin Jin

arXiv:2405.17719·cs.CV·February 21, 2025

Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?

Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song, Sipeng Zheng, Qin Jin

PDF

Open Access 1 Repo

TL;DR

This paper evaluates whether egocentric video-language models truly understand hand-object interactions, introduces a new benchmark EgoHOIBench to test their limitations, and proposes EgoNCE++ to improve their fine-grained understanding and performance.

Contribution

It introduces EgoHOIBench for assessing egocentric models' understanding of hand-object interactions and proposes EgoNCE++ to enhance model performance through improved supervision and contrastive learning.

Findings

01

EgoVLMs struggle with simple modifications in interaction descriptions.

02

EgoNCE++ significantly improves model performance on various tasks.

03

Models show greater difficulty recognizing verbs than nouns in interactions.

Abstract

Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by simple modifications, such as changing the verbs or nouns in interaction descriptions, with models struggling to distinguish between these changes. This raises the question: Do EgoVLMs truly understand hand-object interactions? To address this question, we introduce a benchmark called EgoHOIBench, revealing the performance limitation of current egocentric models when confronted with such challenges. We attribute this performance gap to insufficient fine-grained supervision and the greater difficulty EgoVLMs experience in recognizing verbs compared to nouns. To tackle these issues, we propose a novel asymmetric contrastive objective named EgoNCE++. For the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

xuboshen/egoncepp
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAction Observation and Synchronization · Social and Intergroup Psychology

MethodsFocus