A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

Guohuan Xie; Syed Ariff Syed Hesham; Wenya Guo; Bing Li; Ming-Ming Cheng; Guolei Sun; Yun Liu

arXiv:2506.13552·cs.CV·June 17, 2025

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

Guohuan Xie, Syed Ariff Syed Hesham, Wenya Guo, Bing Li, Ming-Ming Cheng, Guolei Sun, Yun Liu

PDF

Open Access

TL;DR

This survey comprehensively reviews recent advances, challenges, and future prospects in Video Scene Parsing, emphasizing deep learning methods, technical hurdles, and benchmarking standards across various vision tasks.

Contribution

It provides a holistic analysis of VSP evolution, compares datasets and metrics, and discusses emerging trends and research directions in the field.

Findings

01

Deep learning has significantly advanced VSP performance.

02

Maintaining temporal consistency remains a key challenge.

03

Transformer-based architectures are increasingly prominent in VSP.

Abstract

Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a holistic review of recent advances in VSP, covering a wide array of vision tasks, including Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), as well as Video Tracking and Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We systematically analyze the evolution from traditional hand-crafted features to modern deep learning paradigms -- spanning from fully convolutional networks to the latest transformer-based architectures -- and assess their effectiveness in capturing both local and global temporal contexts. Furthermore, our review critically discusses the technical challenges, ranging from…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Image and Video Retrieval Techniques · Multimodal Machine Learning Applications · Video Analysis and Summarization