AssembleNet++: Assembling Modality Representations via Attention   Connections

Michael S. Ryoo; AJ Piergiovanni; Juhana Kangaspunta; Anelia Angelova

arXiv:2008.08072·cs.CV·August 19, 2020

AssembleNet++: Assembling Modality Representations via Attention Connections

Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta, Anelia Angelova

PDF

1 Repo

TL;DR

AssembleNet++ introduces a novel attention-based video model that dynamically integrates semantic object information with appearance and motion features, achieving state-of-the-art activity recognition performance without pre-training.

Contribution

The paper presents AssembleNet++, a new model with peer-attention that learns feature importance across modalities, improving existing architectures for video understanding.

Findings

01

Outperforms previous models on activity recognition datasets

02

Peer-attention effectively learns feature importance dynamically

03

Applicable to various architectures, enhancing their performance

Abstract

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features at each convolutional block of the network. A new network component named peer-attention is introduced, which dynamically learns the attention weights using another block or input modality. Even without pre-training, our models outperform the previous work on standard public activity recognition datasets with continuous videos, establishing new state-of-the-art. We also confirm that our findings of having neural connections from the object modality and the use of peer-attention is generally applicable for different existing architectures, improving their performances. We name our model explicitly as AssembleNet++. The code will be available at:…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

google-research/google-research/tree/master/assemblenet
tf

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsPeer-attention