Weakly-Supervised Action Localization by Generative Attention Modeling

Baifeng Shi; Qi Dai; Yadong Mu; Jingdong Wang

arXiv:2003.12424·cs.CV·March 31, 2020·20 cites

Weakly-Supervised Action Localization by Generative Attention Modeling

Baifeng Shi, Qi Dai, Yadong Mu, Jingdong Wang

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces a novel weakly-supervised action localization method using a conditional Variational Auto-Encoder to better distinguish action frames from context, addressing the common action-context confusion issue.

Contribution

It proposes a probabilistic modeling approach with conditional VAE to improve action localization accuracy under weak supervision, a novel application in this domain.

Findings

01

Outperforms existing methods on THUMOS14 and ActivityNet1.2 datasets.

02

Effectively reduces action-context confusion in localization.

03

Demonstrates the benefit of probabilistic modeling for weakly-supervised learning.

Abstract

Weakly-supervised temporal action localization is a problem of learning an action localization model with only video-level action labeling available. The general framework largely relies on the classification activation, which employs an attention model to identify the action-related frames and then categorizes them into different classes. Such method results in the action-context confusion issue: context frames near action clips tend to be recognized as action frames themselves, since they are closely related to the specific classes. To solve the problem, in this paper we propose to model the class-agnostic frame-wise probability conditioned on the frame attention using conditional Variational Auto-Encoder (VAE). With the observation that the context exhibits notable difference from the action at representation level, a probabilistic model, i.e., conditional VAE, is learned to model…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

bfshi/DGAM-Weakly-Supervised-Action-Localization
pytorchOfficial

Videos

Weakly-Supervised Action Localization by Generative Attention Modeling· youtube

Taxonomy

TopicsHuman Pose and Action Recognition · Anomaly Detection Techniques and Applications · Multimodal Machine Learning Applications

MethodsUSD Coin Customer Service Number +1-833-534-1729