What You Say Is What You Show: Visual Narration Detection in   Instructional Videos

Kumar Ashutosh; Rohit Girdhar; Lorenzo Torresani; Kristen Grauman

arXiv:2301.02307·cs.CV·July 20, 2023·1 cites

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

PDF

Open Access

TL;DR

This paper introduces a novel task of visual narration detection in instructional videos, proposing a multi-modal method that effectively identifies whether narrations are visually depicted, improving video summarization and alignment.

Contribution

The paper presents WYS^2, a new weakly supervised approach leveraging multi-modal cues and pseudo-labeling for visual narration detection in noisy instructional videos.

Findings

01

WYS^2 outperforms strong baselines in visual narration detection.

02

The method improves summarization and temporal alignment of instructional videos.

03

Effective detection achieved with only weakly labeled data.

Abstract

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the narrations do not always describe the actions demonstrated in the video. To address this problem we introduce the novel task of visual narration detection, which entails determining whether a narration is visually depicted by the actions in the video. We propose What You Say is What You Show (WYS^2), a method that leverages multi-modal cues and pseudo-labeling to learn to detect visual narrations with only weakly labeled data. Our model successfully detects visual narrations in in-the-wild videos, outperforming strong baselines, and we demonstrate its impact for state-of-the-art summarization and temporal alignment of instructional videos.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Video Analysis and Summarization