A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction

Jinxi Xiang; Siyu Hou; Yuchen Li; Ryan Quinton; Xiaoming Zhang; Feyisope Eweje; Xiangde Luo; Yijiang Chen; Zhe Li; Colin Bergstrom; Ted Kim; Sierra Willens; Francesca Maria Olguin; Matthew Abikenari; Andrew Heider; Sanjeeth Rajaram; Joel Neal; Maximilian Diehn; Xiang Zhou; Ruijiang Li

arXiv:2604.03630·cs.AI·April 7, 2026

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction

Jinxi Xiang, Siyu Hou, Yuchen Li, Ryan Quinton, Xiaoming Zhang, Feyisope Eweje, Xiangde Luo, Yijiang Chen, Zhe Li, Colin Bergstrom, Ted Kim, Sierra Willens, Francesca Maria Olguin, Matthew Abikenari, Andrew Heider, Sanjeeth Rajaram, Joel Neal, Maximilian Diehn, Xiang Zhou

PDF

TL;DR

STORM is a multimodal foundation model that integrates spatial transcriptomics and histology to improve biological discovery and clinical prediction across multiple organs and platforms.

Contribution

The paper introduces STORM, a hierarchical model trained on extensive data that bridges imaging and omics, enhancing spatial analysis and clinical outcome prediction.

Findings

01

Improves spatial domain discovery with biologically coherent tissue maps.

02

Outperforms existing methods in predicting gene expression from H&E images.

03

Enhances immunotherapy response prediction and prognostication across diverse cohorts.

Abstract

Spatial transcriptomics (ST) enables gene expression mapping within anatomical context but remains costly and low-throughput. Hematoxylin and eosin (H\&E) staining offers rich morphology yet lacks molecular resolution. We present \textbf{\ours} (\textbf{S}patial \textbf{T}ranscriptomics and hist\textbf{O}logy \textbf{R}epresentation \textbf{M}odel), a foundation model trained on 1.2 million spatially resolved transcriptomic profiles with matched histology across 18 organs. Using a hierarchical architecture integrating morphological features, gene expression, and spatial context, STORM bridges imaging and omics through robust molecular--morphological representations. STORM enhances spatial domain discovery, producing biologically coherent tissue maps, and outperforms existing methods in predicting spatial gene expression from H\&E images across 11 tumor types. The model is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.