ExCL: Extractive Clip Localization Using Natural Language Descriptions

Soham Ghosh; Anuva Agarwal; Zarana Parekh; Alexander Hauptmann

arXiv:1904.02755·cs.CL·April 8, 2019·78 cites

ExCL: Extractive Clip Localization Using Natural Language Descriptions

Soham Ghosh, Anuva Agarwal, Zarana Parekh, Alexander Hauptmann

PDF

Open Access 1 Repo

TL;DR

This paper introduces ExCL, a novel extractive method for clip localization in videos based on natural language descriptions, which predicts start and end frames directly through cross-modal interactions, outperforming previous methods.

Contribution

The paper proposes a simple, effective extractive approach that leverages cross-modal interactions to directly predict clip boundaries, eliminating the need for complex proposal and ranking mechanisms.

Findings

01

Significantly outperforms state-of-the-art on two datasets

02

Achieves comparable performance on a third dataset

03

Demonstrates the effectiveness of direct start-end frame prediction

Abstract

The task of retrieving clips within videos based on a given natural language query requires cross-modal reasoning over multiple frames. Prior approaches such as sliding window classifiers are inefficient, while text-clip similarity driven ranking-based approaches such as segment proposal networks are far more complicated. In order to select the most relevant video clip corresponding to the given text description, we propose a novel extractive approach that predicts the start and end frames by leveraging cross-modal interactions between the text and video - this removes the need to retrieve and re-rank multiple proposal segments. Using recurrent networks we encode the two modalities into a joint representation which is then used in different variants of start-end frame predictor networks. Through extensive experimentation and ablative analysis, we demonstrate that our simple and elegant…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

jayleicn/TVRetrieval
pytorch

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsVideo Analysis and Summarization · Multimodal Machine Learning Applications · Human Pose and Action Recognition