Video Playback Rate Perception for Self-supervisedSpatio-Temporal   Representation Learning

Yuan Yao; Chang Liu; Dezhao Luo; Yu Zhou; Qixiang Ye

arXiv:2006.11476·cs.CV·June 23, 2020·1 cites

Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning

Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, Qixiang Ye

PDF

Open Access 1 Repo

TL;DR

This paper introduces a novel self-supervised learning method called Video Playback Rate Perception (PRP) that enhances spatio-temporal video representations by leveraging playback rate classification and reconstruction, improving performance on action recognition and retrieval tasks.

Contribution

The paper proposes a new self-supervised approach, PRP, combining discriminative and generative models to better capture multi-scale temporal features in videos.

Findings

01

PRP outperforms existing self-supervised models on key video tasks.

02

The method effectively captures both long-term and short-term temporal features.

03

Experimental results demonstrate significant performance improvements.

Abstract

In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this paper, we propose a novel self-supervised method, referred to as video Playback Rate Perception (PRP), to learn spatio-temporal representation in a simple-yet-effective way. PRP roots in a dilated sampling strategy, which produces self-supervision signals about video playback rates for representation model learning. PRP is implemented with a feature encoder, a classification module, and a reconstructing decoder, to achieve spatio-temporal semantic retention in a collaborative discrimination-generation manner. The discriminative perception model follows a feature encoder to prefer perceiving low temporal resolution and long-term representation by classifying fast-forward…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yuanyao366/PRP
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · Advanced Vision and Imaging · Multimodal Machine Learning Applications