Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

Seunghwan Bang; Hwanjun Song

arXiv:2603.13091·cs.CV·April 20, 2026

Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

Seunghwan Bang, Hwanjun Song

PDF

TL;DR

This paper introduces VAEX-BENCH, a new benchmark for evaluating multimodal large language models' ability to perform complex, abstractive spatiotemporal reasoning over videos, highlighting current limitations.

Contribution

It formalizes abstractive spatiotemporal reasoning, creates a synthetic dataset, and provides a comprehensive evaluation framework for MLLMs on these challenging tasks.

Findings

01

State-of-the-art MLLMs struggle with abstractive reasoning tasks.

02

The benchmark reveals specific limitations and bottlenecks in current models.

03

Extensive experiments compare extractive and abstractive reasoning capabilities.

Abstract

The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extractive reasoning, where answers can be explicitly presented within spatiotemporal events. It remains unclear whether multimodal large language models can instead perform abstractive spatiotemporal reasoning, which requires integrating observations over time, combining dispersed cues, and inferring implicit spatial and contextual structure. To address this gap, we formalize abstractive spatiotemporal reasoning from videos by introducing a structured evaluation taxonomy that systematically targets its core dimensions and constructs a controllable, scenario-driven synthetic egocentric video dataset tailored to evaluate abstractive spatiotemporal reasoning capabilities, spanning object-, room-, and floor-plan-level scenarios. Based on this…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.