When Large Vision-Language Models Meet Person Re-Identification

Qizao Wang; Bin Li; Xiangyang Xue

arXiv:2411.18111·cs.CV·May 12, 2026

When Large Vision-Language Models Meet Person Re-Identification

Qizao Wang, Bin Li, Xiangyang Xue

PDF

TL;DR

This paper introduces LVLM-ReID, a framework that leverages large vision-language models to improve person re-identification by generating and refining semantic tokens representing identity features.

Contribution

It proposes a novel method to utilize LVLMs for ReID by generating semantic tokens guided by instructions and refined through a Semantic-Guided Interaction module.

Findings

01

Achieves competitive results on multiple ReID benchmarks.

02

Operates without additional image-text annotations.

03

Effectively captures rich semantic cues for identity recognition.

Abstract

Large Vision-Language Models (LVLMs) that incorporate visual models and large language models have achieved impressive results across cross-modal understanding and reasoning tasks. In recent years, person re-identification (ReID) has also started to explore cross-modal semantics to improve the accuracy of identity recognition. However, effectively utilizing LVLMs for ReID remains an open challenge. While LVLMs operate under a generative paradigm by predicting the next output word, ReID requires the extraction of discriminative identity features to match pedestrians across cameras. In this paper, we propose LVLM-ReID, a novel framework that harnesses the strengths of LVLMs to promote ReID. Specifically, we employ instructions to guide the LVLM in generating one semantic token that encapsulates key appearance semantics from the person image. This token is further refined through our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.