TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Fanheng Kong; Jingyuan Zhang; Hongzhi Zhang; Shi Feng; Daling Wang; Linhao Yu; Xingguang Ji; Yu Tian; Victoria W.; Fuzheng Zhang

arXiv:2505.20124·cs.CV·June 10, 2025

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang, Shi Feng, Daling Wang, Linhao Yu, Xingguang Ji, Yu Tian, Victoria W., Fuzheng Zhang

PDF

Open Access 1 Repo 1 Datasets 1 Video

TL;DR

TUNA is a new benchmark designed to evaluate fine-grained temporal understanding in dense dynamic videos through captioning and QA tasks, highlighting key challenges and guiding future improvements.

Contribution

It introduces a comprehensive, multi-faceted benchmark for dense dynamic videos, addressing the holistic temporal understanding often overlooked in existing datasets.

Findings

01

Models struggle with detailed action descriptions.

02

Multi-subject understanding remains limited.

03

Camera motion sensitivity is inadequate.

Abstract

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often treat these properties separately or narrowly focus on specific aspects, overlooking the holistic nature of video content. To address this, we introduce TUNA, a temporal-oriented benchmark for fine-grained understanding on dense dynamic videos, with two complementary tasks: captioning and QA. Our TUNA features diverse video scenarios and dynamics, assisted by interpretable and robust evaluation criteria. We evaluate several leading models on our benchmark, providing fine-grained performance assessments across various dimensions. This evaluation reveals key challenges in video temporal understanding, such as limited action description, inadequate multi-subject…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

friedrichor/TUNA
pytorchOfficial

Datasets

friedrichor/TUNA-Bench
dataset· 118 dl
118 dl

Videos

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos· underline

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Generative Adversarial Networks and Image Synthesis

MethodsFocus