Open-vocabulary Video Question Answering: A New Benchmark for Evaluating   the Generalizability of Video Question Answering Models

Dohwan Ko; Ji Soo Lee; Miso Choi; Jaewon Chu; Jihwan Park; Hyunwoo J.; Kim

arXiv:2308.09363·cs.CV·August 21, 2023·2 cites

Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models

Dohwan Ko, Ji Soo Lee, Miso Choi, Jaewon Chu, Jihwan Park, Hyunwoo J., Kim

PDF

Open Access 1 Repo

TL;DR

This paper introduces a new benchmark, OVQA, for evaluating the ability of VideoQA models to generalize to rare and unseen answers, and proposes a GNN-based soft verbalizer to enhance model performance on these answers.

Contribution

The paper presents the OVQA benchmark for open-vocabulary VideoQA and introduces a GNN-based soft verbalizer to improve generalization to rare and unseen answers.

Findings

01

The GNN-based soft verbalizer improves answer prediction accuracy.

02

Models trained with OVQA generalize better to out-of-vocabulary answers.

03

Benchmark results highlight the gap in current models' ability to handle rare answers.

Abstract

Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, the majority of previous VideoQA models formulate open-ended VideoQA as a classification task to classify the video-question pairs into a fixed answer set, i.e., closed-vocabulary, which contains only frequent answers (e.g., top-1000 answers). This leads the model to be biased toward only frequent answers and fail to generalize on out-of-vocabulary answers. We hence propose a new benchmark, Open-vocabulary Video Question Answering (OVQA), to measure the generalizability of VideoQA models by considering rare and unseen answers. In addition, in order to improve the model's generalization power,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

mlvlab/ovqa
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Topic Modeling · Domain Adaptation and Few-Shot Learning

Methodsfail