VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Jianshu Zhang; Dongyu Yao; Renjie Pi; Paul Pu Liang; Yi R. Fung

arXiv:2502.12084·cs.CL·July 3, 2025

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Jianshu Zhang, Dongyu Yao, Renjie Pi, Paul Pu Liang, Yi R. Fung

PDF

Open Access 1 Repo 1 Datasets 1 Video

TL;DR

VLM2-Bench evaluates vision-language models' ability to link matching visual cues, revealing significant performance gaps and guiding future improvements in visual reasoning and training paradigms.

Contribution

Introduces VLM2-Bench, a comprehensive benchmark for assessing visual cue linking in VLMs, and provides insights into their limitations and directions for enhancement.

Findings

01

Models struggle with linking visual cues accurately.

02

Performance varies significantly across different models and prompting methods.

03

Recommendations for improving visual reasoning and training strategies.

Abstract

Visually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even without knowing who they are. Despite the extensive knowledge that vision-language models (VLMs) possess, it remains largely unexplored whether they are capable of performing this fundamental task. To address this, we introduce \textbf{VLM2-Bench}, a benchmark designed to assess whether VLMs can Visually Link Matching cues, with 9 subtasks and over 3,000 test cases. Comprehensive evaluation across twelve VLMs, along with further analysis of various language-side and vision-side prompting methods, leads to a total of eight key findings. We identify critical challenges in models' ability to link visual cues, highlighting a significant performance gap. Based on these insights, we advocate for (i) enhancing core visual capabilities to improve…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

vlm2-bench/VLM2-Bench
pytorch

Datasets

Sterzhang/vlm2-bench
dataset· 65 dl
65 dl

Videos

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues· underline

Taxonomy

TopicsRetinal Imaging and Analysis