Anatomical Attention Alignment representation for Radiology Report Generation

Quang Vinh Nguyen; Minh Duc Nguyen; Thanh Hoang Son Vo; Hyung-Jeong Yang; Soo-Hyung Kim

arXiv:2505.07689·cs.CV·May 13, 2025

Anatomical Attention Alignment representation for Radiology Report Generation

Quang Vinh Nguyen, Minh Duc Nguyen, Thanh Hoang Son Vo, Hyung-Jeong Yang, Soo-Hyung Kim

PDF

Open Access 1 Repo

TL;DR

This paper introduces A3Net, a novel framework that enhances radiology report generation by integrating anatomical knowledge with visual features, leading to more accurate and interpretable reports.

Contribution

A3Net is the first model to incorporate a knowledge dictionary of anatomical structures for improved visual-textual alignment in radiology report generation.

Findings

01

Significant improvement in report accuracy on IU X-Ray and MIMIC-CXR datasets.

02

Enhanced interpretability and semantic reasoning in generated reports.

03

Better cross-modal alignment between image regions and anatomical entities.

Abstract

Automated Radiology report generation (RRG) aims at producing detailed descriptions of medical images, reducing radiologists' workload and improving access to high-quality diagnostic services. Existing encoder-decoder models only rely on visual features extracted from raw input images, which can limit the understanding of spatial structures and semantic relationships, often resulting in suboptimal text generation. To address this, we propose Anatomical Attention Alignment Network (A3Net), a framework that enhance visual-textual understanding by constructing hyper-visual representations. Our approach integrates a knowledge dictionary of anatomical structures with patch-level visual features, enabling the model to effectively associate image regions with their corresponding anatomical entities. This structured representation improves semantic reasoning, interpretability, and cross-modal…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

vinh-ai/a3net
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Topic Modeling · Generative Adversarial Networks and Image Synthesis

MethodsSoftmax · Attention Is All You Need