How to Understand Masked Autoencoders

Shuhao Cao; Peng Xu; David A. Clifton

arXiv:2202.03670·cs.CV·February 10, 2022·23 cites

How to Understand Masked Autoencoders

Shuhao Cao, Peng Xu, David A. Clifton

PDF

Open Access

TL;DR

This paper provides the first unified theoretical framework to explain the high expressivity and success of Masked Autoencoders (MAE) in vision learning, bridging the gap with linguistic masked autoencoding.

Contribution

It introduces a mathematical understanding of MAE's patch-based attention through integral kernels and operator theory, offering new insights into its effectiveness.

Findings

01

Provides a theoretical explanation for MAE's expressivity

02

Uses integral kernel and operator theory to analyze attention mechanisms

03

Answers key questions about MAE's success with mathematical rigor

Abstract

"Masked Autoencoders (MAE) Are Scalable Vision Learners" revolutionizes the self-supervised learning method in that it not only achieves the state-of-the-art for image pre-training, but is also a milestone that bridges the gap between visual and linguistic masked autoencoding (BERT-style) pre-trainings. However, to our knowledge, to date there are no theoretical perspectives to explain the powerful expressivity of MAE. In this paper, we, for the first time, propose a unified theoretical framework that provides a mathematical understanding for MAE. Specifically, we explain the patch-based attention approaches of MAE using an integral kernel under a non-overlapping domain decomposition setting. To help the research community to further comprehend the main reasons of the great success of MAE, based on our framework, we pose five questions and answer them with mathematical rigor using…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Generative Adversarial Networks and Image Synthesis · Multimodal Machine Learning Applications

MethodsMasked autoencoder