SeMask: Semantically Masked Transformers for Semantic Segmentation

Jitesh Jain; Anukriti Singh; Nikita Orlov; Zilong Huang; Jiachen Li,; Steven Walton; Humphrey Shi

arXiv:2112.12782·cs.CV·April 14, 2022·31 cites

SeMask: Semantically Masked Transformers for Semantic Segmentation

Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li,, Steven Walton, Humphrey Shi

PDF

Open Access 1 Repo

TL;DR

SeMask introduces a semantic attention framework that enhances hierarchical transformer encoders for semantic segmentation by incorporating semantic priors, leading to state-of-the-art results on ADE20K and Cityscapes datasets.

Contribution

The paper proposes SeMask, a novel semantic attention method that integrates semantic priors into transformer encoders during finetuning, improving segmentation performance.

Findings

01

Achieved 58.25% mIoU on ADE20K, setting a new state-of-the-art.

02

Improved Cityscapes mIoU by over 3%.

03

Enhanced encoder performance with minimal increase in FLOPs.

Abstract

Finetuning a pretrained backbone in the encoder part of an image transformer network has been the traditional approach for the semantic segmentation task. However, such an approach leaves out the semantic context that an image provides during the encoding stage. This paper argues that incorporating semantic information of the image into pretrained hierarchical transformer-based backbones while finetuning improves the performance considerably. To achieve this, we propose SeMask, a simple and effective framework that incorporates semantic information into the encoder with the help of a semantic attention operation. In addition, we use a lightweight semantic decoder during training to provide supervision to the intermediate semantic prior maps at every stage. Our experiments demonstrate that incorporating semantic priors enhances the performance of the established hierarchical encoders…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Picsart-AI-Research/SeMask-Segmentation
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Neural Network Applications · Multimodal Machine Learning Applications · Domain Adaptation and Few-Shot Learning

MethodsMulti-Head Attention · Attention Is All You Need · Linear Layer · Byte Pair Encoding · Position-Wise Feed-Forward Layer · Label Smoothing · Dropout · Layer Normalization · Adam · Absolute Position Encodings