Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography

Simon B\"ohi; Irene Cannistraci; Sergio Mu\~noz Gonzalez; Moritz Vandenhirtz; Sonia Laguna; Samuel Ruiperez-Campillo; Max Kr\"ahenmann; Andrea Agostini; Ece Ozkan; Thomas M. Sutter; Julia E. Vogt

arXiv:2604.15096·cs.CV·April 17, 2026

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography

Simon B\"ohi, Irene Cannistraci, Sergio Mu\~noz Gonzalez, Moritz Vandenhirtz, Sonia Laguna, Samuel Ruiperez-Campillo, Max Kr\"ahenmann, Andrea Agostini, Ece Ozkan, Thomas M. Sutter, Julia E. Vogt

PDF

TL;DR

This paper introduces LAMAE, a multi-view masked autoencoder for echocardiography that captures cross-view information in latent space, improving cardiac representation and transferability across patient groups.

Contribution

LAMAE is the first model to incorporate multi-view latent attention in masked autoencoders for echocardiography, enabling holistic cardiac analysis from partial, multi-view data.

Findings

01

LAMAE effectively reconstructs cardiac function from partial multi-view observations.

02

Pretraining on MIMIC-IV-ECHO improves ICD-10 code prediction accuracy.

03

Transfer learning from adult to pediatric data remains effective despite anatomical differences.

Abstract

Echocardiography is a widely used modality for cardiac assessment due to its non-invasive and cost-effective nature, but the sparse and heterogeneous spatiotemporal views of the heart pose distinct challenges. Existing masked autoencoder (MAE) approaches typically process images or short clips independently, failing to capture the inherent multi-view structure required for coherent cardiac representation. We introduce Latent Attention Masked Autoencoder (LAMAE), a foundation model architecture tailored to the multi-view nature of medical imaging. LAMAE augments the standard MAE with a latent attention module that enables information exchange across frames and views directly in latent space. This allows the model to aggregate variable-length sequences and distinct views, reconstructing a holistic representation of cardiac function from partial observations. We pretrain LAMAE on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.