Online Action Detection in Streaming Videos with Time Buffers

Bowen Zhang; Hao Chen; Meng Wang; Yuanjun Xiong

arXiv:2010.03016·cs.CV·October 8, 2020·1 cites

Online Action Detection in Streaming Videos with Time Buffers

Bowen Zhang, Hao Chen, Meng Wang, Yuanjun Xiong

PDF

Open Access

TL;DR

This paper introduces a new online action detection framework that accounts for broadcast delay in streaming videos, improving detection accuracy by utilizing a small buffer time rather than immediate prediction.

Contribution

It proposes a novel problem setting for online action detection that incorporates buffer time, along with a new detection framework tailored for this setting.

Findings

01

Significant accuracy improvements over existing models.

02

Effective use of buffer time enhances detection performance.

03

Validated on three standard benchmarks.

Abstract

We formulate the problem of online temporal action detection in live streaming videos, acknowledging one important property of live streaming videos that there is normally a broadcast delay between the latest captured frame and the actual frame viewed by the audience. The standard setting of the online action detection task requires immediate prediction after a new frame is captured. We illustrate that its lack of consideration of the delay is imposing unnecessary constraints on the models and thus not suitable for this problem. We propose to adopt the problem setting that allows models to make use of the small `buffer time' incurred by the delay in live streaming videos. We design an action start and end detection framework for this online with buffer setting with two major components: flattened I3D and window-based suppression. Experiments on three standard temporal action detection…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · Video Analysis and Summarization · Video Coding and Compression Technologies