SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding

Jiwook Han; Geo Ahn; Youngrae Kim; Jinwoo Choi

arXiv:2603.25733·cs.CV·March 27, 2026

SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding

Jiwook Han, Geo Ahn, Youngrae Kim, Jinwoo Choi

PDF

Open Access

TL;DR

SlotVTG enhances video temporal grounding by integrating object-centric reasoning into multimodal models, significantly improving out-of-domain generalization with minimal additional training.

Contribution

It introduces a lightweight slot adapter that promotes object-centric visual reasoning in MLLMs, improving OOD robustness without extensive retraining.

Findings

01

Significant improvement in OOD generalization on VTG benchmarks.

02

Maintains competitive In-Domain performance.

03

Minimal computational overhead compared to full retraining.

Abstract

Multimodal Large Language Models (MLLMs) have shown strong performance on Video Temporal Grounding (VTG). However, their coarse recognition capabilities are insufficient for fine-grained temporal understanding, making task-specific fine-tuning indispensable. This fine-tuning causes models to memorize dataset-specific shortcuts rather than faithfully grounding in the actual visual content, leading to poor Out-of-Domain (OOD) generalization. Object-centric learning offers a promising remedy by decomposing scenes into entity-level representations, but existing approaches require re-running the entire multi-stage training pipeline from scratch. We propose SlotVTG, a framework that steers MLLMs toward object-centric, input-grounded visual reasoning at minimal cost. SlotVTG introduces a lightweight slot adapter that decomposes visual tokens into abstract slots via slot attention and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Domain Adaptation and Few-Shot Learning · Generative Adversarial Networks and Image Synthesis