Rethinking Attention Gated with Hybrid Dual Pyramid Transformer-CNN for   Generalized Segmentation in Medical Imaging

Fares Bougourzi; Fadi Dornaika; Abdelmalik Taleb-Ahmed; Vinh Truong; Hoang

arXiv:2404.18199·eess.IV·April 30, 2024

Rethinking Attention Gated with Hybrid Dual Pyramid Transformer-CNN for Generalized Segmentation in Medical Imaging

Fares Bougourzi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Vinh Truong, Hoang

PDF

Open Access 1 Repo

TL;DR

This paper introduces a novel hybrid CNN-Transformer architecture with attention gates and pyramid input for improved medical image segmentation, demonstrating state-of-the-art results across multiple tasks.

Contribution

The paper presents a new hybrid encoder architecture combining CNN and Transformer with dual attention gates and pyramid inputs for enhanced segmentation performance.

Findings

01

Achieved state-of-the-art results on multiple medical segmentation tasks.

02

Demonstrated strong generalization across different datasets.

03

Efficiently captures multi-scale features and long-range dependencies.

Abstract

Inspired by the success of Transformers in Computer vision, Transformers have been widely investigated for medical imaging segmentation. However, most of Transformer architecture are using the recent transformer architectures as encoder or as parallel encoder with the CNN encoder. In this paper, we introduce a novel hybrid CNN-Transformer segmentation architecture (PAG-TransYnet) designed for efficiently building a strong CNN-Transformer encoder. Our approach exploits attention gates within a Dual Pyramid hybrid encoder. The contributions of this methodology can be summarized into three key aspects: (i) the utilization of Pyramid input for highlighting the prominent features at different scales, (ii) the incorporation of a PVT transformer to capture long-range dependencies across various resolutions, and (iii) the implementation of a Dual-Attention Gate mechanism for effectively fusing…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

faresbougourzi/pagtransynet
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBrain Tumor Detection and Classification

MethodsAttention Is All You Need · Dropout · Residual Connection · Softmax · Spatial-Reduction Attention · Position-Wise Feed-Forward Layer · Byte Pair Encoding · Absolute Position Encodings · Pyramid Vision Transformer · Linear Layer