Minimum Data, Maximum Impact: 20 annotated samples for explainable lung nodule classification

Luisa Gall\'ee; Catharina Silvia Lisson; Christoph Gerhard Lisson; Daniela Drees; Felix Weig; Daniel Vogele; Meinrad Beer; Michael G\"otz

arXiv:2508.00639·cs.CV·August 4, 2025

Minimum Data, Maximum Impact: 20 annotated samples for explainable lung nodule classification

Luisa Gall\'ee, Catharina Silvia Lisson, Christoph Gerhard Lisson, Daniela Drees, Felix Weig, Daniel Vogele, Meinrad Beer, Michael G\"otz

PDF

Open Access

TL;DR

This paper demonstrates that using a small set of 20 annotated lung nodule images to train a generative model can produce synthetic data that significantly improves the performance of explainable classification models in medical imaging.

Contribution

The study introduces a method to synthesize attribute-annotated medical images using a diffusion model trained on only 20 samples, enhancing explainable lung nodule classification.

Findings

01

Attribute prediction accuracy increased by 13.4%.

02

Target prediction accuracy increased by 1.8%.

03

Synthetic data effectively compensates for limited real annotations.

Abstract

Classification models that provide human-interpretable explanations enhance clinicians' trust and usability in medical image diagnosis. One research focus is the integration and prediction of pathology-related visual attributes used by radiologists alongside the diagnosis, aligning AI decision-making with clinical reasoning. Radiologists use attributes like shape and texture as established diagnostic criteria and mirroring these in AI decision-making both enhances transparency and enables explicit validation of model outputs. However, the adoption of such models is limited by the scarcity of large-scale medical image datasets annotated with these attributes. To address this challenge, we propose synthesizing attribute-annotated data using a generative model. We enhance the Diffusion Model with attribute conditioning and train it using only 20 attribute-labeled lung nodule samples from…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsExplainable Artificial Intelligence (XAI) · AI in cancer detection · COVID-19 diagnosis using AI