NAVERO: Unlocking Fine-Grained Semantics for Video-Language   Compositionality

Chaofan Tao; Gukyeong Kwon; Varad Gunjal; Hao Yang; Zhaowei Cai,; Yonatan Dukler; Ashwin Swaminathan; R. Manmatha; Colin Jon Taylor; Stefano; Soatto

arXiv:2408.09511·cs.CV·August 20, 2024

NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality

Chaofan Tao, Gukyeong Kwon, Varad Gunjal, Hao Yang, Zhaowei Cai,, Yonatan Dukler, Ashwin Swaminathan, R. Manmatha, Colin Jon Taylor, Stefano, Soatto

PDF

Open Access

TL;DR

This paper introduces NAVERO, a training method that enhances video-language models' ability to understand complex object-attribute-action compositions over time, by using negative text augmentation and a specialized loss, improving compositional understanding.

Contribution

The paper presents NAVERO, a novel training approach with negative-augmented loss and a new benchmark AARO for evaluating fine-grained video-language compositionality, addressing temporal relation challenges.

Findings

01

NAVERO significantly outperforms existing methods in compositional understanding.

02

NAVERO maintains strong performance on traditional video-text retrieval tasks.

03

The AARO benchmark effectively evaluates action-based compositional understanding.

Abstract

We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes particularly challenging for video data since the compositional relations rapidly change over time in videos. We first build a benchmark named AARO to evaluate composition understanding related to actions on top of spatial concepts. The benchmark is constructed by generating negative texts with incorrect action descriptions for a given video and the model is expected to pair a positive text with its corresponding video. Furthermore, we propose a training method called NAVERO which utilizes video-text data augmented with negative texts to enhance composition understanding. We also develop a negative-augmented visual-language matching loss which is used explicitly to benefit from the generated negative text. We…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications