Transferring Domain-Agnostic Knowledge in Video Question Answering

Tianran Wu; Noa Garcia; Mayu Otani; Chenhui Chu; Yuta Nakashima and; Haruo Takemura

arXiv:2110.13395·cs.CV·October 27, 2021·5 cites

Transferring Domain-Agnostic Knowledge in Video Question Answering

Tianran Wu, Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima and, Haruo Takemura

PDF

Open Access

TL;DR

This paper introduces a transfer learning approach for VideoQA that leverages domain-agnostic knowledge to improve performance, supported by a new dataset and experimental validation.

Contribution

It proposes a novel transfer learning framework using domain-agnostic knowledge and creates a new large-scale VideoQA dataset for evaluation.

Findings

01

Domain-agnostic knowledge is transferable across tasks.

02

The proposed framework significantly improves VideoQA accuracy.

Abstract

Video question answering (VideoQA) is designed to answer a given question based on a relevant video clip. The current available large-scale datasets have made it possible to formulate VideoQA as the joint understanding of visual and language information. However, this training procedure is costly and still less competent with human performance. In this paper, we investigate a transfer learning method by the introduction of domain-agnostic knowledge and domain-specific knowledge. First, we develop a novel transfer learning framework, which finetunes the pre-trained model by applying domain-agnostic knowledge as the medium. Second, we construct a new VideoQA dataset with 21,412 human-generated question-answer samples for comparable transfer of knowledge. Our experiments show that: (i) domain-agnostic knowledge is transferable and (ii) our proposed transfer learning framework can boost…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Domain Adaptation and Few-Shot Learning · Human Pose and Action Recognition