Visual Semantic Description Generation with MLLMs for Image-Text Matching

Junyu Chen; Yihua Gao; Mingyong Li

arXiv:2507.08590·cs.MM·July 14, 2025

Visual Semantic Description Generation with MLLMs for Image-Text Matching

Junyu Chen, Yihua Gao, Mingyong Li

PDF

TL;DR

This paper introduces a novel framework using multimodal large language models to generate visual semantic descriptions, significantly improving image-text matching performance and zero-shot generalization across domains.

Contribution

It proposes a new method that leverages MLLMs for semantic parsing to enhance cross-modal alignment in image-text matching tasks.

Findings

01

Significant performance improvements on Flickr30K and MSCOCO datasets.

02

Effective zero-shot generalization to cross-domain image-text matching tasks.

03

Seamless integration with existing ITM models enhances their capabilities.

Abstract

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We propose a novel framework that bridges the modality gap by leveraging multimodal large language models (MLLMs) as visual semantic parsers. By generating rich Visual Semantic Descriptions (VSD), MLLMs provide semantic anchor that facilitate cross-modal alignment. Our approach combines: (1) Instance-level alignment by fusing visual features with VSD to enhance the linguistic expressiveness of image representations, and (2) Prototype-level alignment through VSD clustering to ensure category-level consistency. These modules can be seamlessly integrated into existing ITM models. Extensive experiments on Flickr30K and MSCOCO demonstrate substantial…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.