Support-set bottlenecks for video-text representation learning

Mandela Patrick; Po-Yao Huang; Yuki Asano; Florian Metze; Alexander; Hauptmann; Jo\~ao Henriques; Andrea Vedaldi

arXiv:2010.02824·cs.CV·January 15, 2021·37 cites

Support-set bottlenecks for video-text representation learning

Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander, Hauptmann, Jo\~ao Henriques, Andrea Vedaldi

PDF

Open Access 1 Video

TL;DR

This paper introduces a novel approach to video-text representation learning that relaxes the strict dissimilarity enforcement of contrastive learning by using a generative model to promote semantic sharing among related samples, improving retrieval performance.

Contribution

It proposes a generative-based method that encourages semantic similarity among related samples, addressing limitations of traditional contrastive learning in video-text tasks.

Findings

01

Outperforms existing methods on MSR-VTT, VATEX, ActivityNet, and MSVD datasets.

02

Achieves significant improvements in video-to-text and text-to-video retrieval tasks.

03

Promotes more generalizable and semantically rich representations.

Abstract

The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to be related, such as text and video from the same sample, and pushes away the representations of all other pairs. We posit that this last behaviour is too strict, enforcing dissimilar representations even for samples that are semantically-related -- for example, visually similar videos or ones that share the same depicted action. In this paper, we propose a novel method that alleviates this by leveraging a generative model to naturally push these related samples together: each sample's caption must be reconstructed as a weighted combination of other support samples' visual representations. This simple idea ensures that representations are not overly-specialized to individual samples, are reusable across the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Support-set bottlenecks for video-text representation learning· slideslive

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Domain Adaptation and Few-Shot Learning

MethodsContrastive Learning