Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

Xiaomin Yu; Yi Xin; Yuhui Zhang; Wenjie Zhang; Chonghan Liu; Hanzhen Zhao; Chen Liu; Xiaoxing Hu; Ziyue Qiao; Hao Tang; Xiaobin Hu; Chengwei Qin; Hui Xiong; Yu Qiao; Shuicheng Yan

arXiv:2602.07026·cs.CV·May 11, 2026

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan

PDF

1 Repo

TL;DR

This paper introduces a geometric model of the modality gap in multimodal models and proposes ReAlign and ReVision strategies for efficient, scalable alignment using unpaired data, reducing reliance on costly image-text pairs.

Contribution

It offers a precise geometric characterization of the modality gap and introduces ReAlign and ReVision for effective, training-free alignment and scalable model training with unpaired data.

Findings

01

ReAlign effectively aligns text and image representations using unpaired data.

02

ReVision enables scalable training of multimodal models without large image-text datasets.

03

The proposed methods improve model alignment and scaling efficiency.

Abstract

Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions. Prior approaches to bridge this gap are largely limited by oversimplified isotropic assumptions, hindering their application in large-scale scenarios. In this paper, we address these limitations by precisely characterizing the geometric shape of the modality gap and leveraging it for efficient model scaling. First, we propose the Fixed-frame Modality Gap Theory, which decomposes the modality gap within a frozen reference frame into stable biases and anisotropic residuals. Guided by this precise modeling, we introduce ReAlign, a training-free modality alignment strategy. Utilizing statistics from massive unpaired data,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yu-xm/ReVision
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.