A Discriminative CNN Video Representation for Event Detection

Zhongwen Xu; Yi Yang; Alexander G. Hauptmann

arXiv:1411.4006·cs.CV·November 17, 2014·29 cites

A Discriminative CNN Video Representation for Event Detection

Zhongwen Xu, Yi Yang, Alexander G. Hauptmann

PDF

Open Access

TL;DR

This paper introduces a novel CNN-based video representation using latent concept descriptors and advanced encoding, significantly improving event detection accuracy on large datasets with limited hardware.

Contribution

It proposes a new encoding method and latent concept descriptors for CNN features, achieving state-of-the-art event detection performance.

Findings

01

Improved mAP from 27.6% to 36.8% on TRECVID MEDTest 14

02

Enhanced performance over Dense Trajectories

03

Achieved top results in TRECVID MED 2014 competition

Abstract

In this paper, we propose a discriminative video representation for event detection over a large scale video dataset when only limited hardware resources are available. The focus of this paper is to effectively leverage deep Convolutional Neural Networks (CNNs) to advance event detection, where only frame level static descriptors can be extracted by the existing CNN toolkit. This paper makes two contributions to the inference of CNN video representation. First, while average pooling and max pooling have long been the standard approaches to aggregating frame level static features, we show that performance can be significantly improved by taking advantage of an appropriate encoding method. Second, we propose using a set of latent concept descriptors as the frame descriptor, which enriches visual information while keeping it computationally affordable. The integration of the two…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · Advanced Image and Video Retrieval Techniques · Video Surveillance and Tracking Methods

MethodsAverage Pooling · Max Pooling