Spatiotemporal CNN for Video Object Segmentation

Kai Xu; Longyin Wen; Guorong Li; Liefeng Bo; Qingming Huang

arXiv:1904.02363·cs.CV·April 5, 2019·6 cites

Spatiotemporal CNN for Video Object Segmentation

Kai Xu, Longyin Wen, Guorong Li, Liefeng Bo, Qingming Huang

PDF

Open Access 1 Repo

TL;DR

This paper introduces a unified spatiotemporal CNN model for video object segmentation that leverages adversarial training and a coarse-to-fine attention mechanism to improve segmentation accuracy on challenging datasets.

Contribution

The novel end-to-end trainable model combines a temporal coherence branch with a spatial segmentation branch, utilizing adversarial pretraining and multi-scale attention for enhanced VOS performance.

Findings

01

Achieves state-of-the-art results on DAVIS and Youtube-Object datasets.

02

Effectively captures dynamic appearance and motion cues.

03

Improves segmentation accuracy with a coarse-to-fine attention approach.

Abstract

In this paper, we present a unified, end-to-end trainable spatiotemporal CNN model for VOS, which consists of two branches, i.e., the temporal coherence branch and the spatial segmentation branch. Specifically, the temporal coherence branch pretrained in an adversarial fashion from unlabeled video data, is designed to capture the dynamic appearance and motion cues of video sequences to guide object segmentation. The spatial segmentation branch focuses on segmenting objects accurately based on the learned appearance and motion cues. To obtain accurate segmentation results, we design a coarse-to-fine process to sequentially apply a designed attention module on multi-scale feature maps, and concatenate them to produce the final prediction. In this way, the spatial segmentation branch is enforced to gradually concentrate on object regions. These two branches are jointly fine-tuned on video…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

longyin880815/STCNN
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Neural Network Applications · Visual Attention and Saliency Detection · Video Surveillance and Tracking Methods