Learning Self-Similarity in Space and Time as Generalized Motion for   Video Action Recognition

Heeseung Kwon; Manjin Kim; Suha Kwak; Minsu Cho

arXiv:2102.07092·cs.CV·November 3, 2021

Learning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition

Heeseung Kwon, Manjin Kim, Suha Kwak, Minsu Cho

PDF

Open Access 1 Repo

TL;DR

This paper introduces a novel spatio-temporal self-similarity (STSS) based motion representation for video action recognition, improving the modeling of motion dynamics and outperforming previous methods on multiple benchmarks.

Contribution

The paper proposes the SELFY neural block that captures long-term and fast motions using STSS, which can be integrated into existing architectures for end-to-end training.

Findings

01

Achieves state-of-the-art results on Something-Something-V1 & V2, Diving-48, and FineGym datasets.

02

Effectively models long-term interactions and fast motions in videos.

03

Demonstrates superiority and complementarity over previous motion modeling methods.

Abstract

Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion representation based on spatio-temporal self-similarity (STSS). Given a sequence of frames, STSS represents each local region as similarities to its neighbors in space and time. By converting appearance features into relational values, it enables the learner to better recognize structural patterns in space and time. We leverage the whole volume of STSS and let our model learn to extract an effective motion representation from it. The proposed neural block, dubbed SELFY, can be easily inserted into neural architectures and trained end-to-end without additional supervision. With a sufficient volume of the neighborhood in space and time, it effectively captures long-term…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

arunos728/SELFY
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · Gait Recognition and Analysis · Anomaly Detection Techniques and Applications

MethodsConvolution