VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation

Divake Kumar; Sina Tayebati; Devashri Naik; Ranganath Krishnan; Amit Ranjan Trivedi

arXiv:2604.25235·cs.LG·April 30, 2026

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation

Divake Kumar, Sina Tayebati, Devashri Naik, Ranganath Krishnan, Amit Ranjan Trivedi

PDF

1 Repo

TL;DR

This paper analyzes the uncertainty in vision-language model judges for multimodal evaluation, revealing task-dependent reliability and proposing conformal prediction to calibrate score intervals without retraining.

Contribution

It introduces the first systematic application of conformal prediction to VLM judges, mapping evaluation uncertainty across tasks and identifying factors affecting interval width.

Findings

01

Evaluation uncertainty varies significantly across tasks.

02

Standard metrics fail to capture judge reliability and uncertainty.

03

Interval width correlates with task difficulty and annotation quality.

Abstract

Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability. We study this problem through conformal prediction, a distribution-free framework that converts a judge's point score into a calibrated prediction interval using only score-token log-probabilities, with no retraining. We present the first systematic analysis of conformal prediction for VLM-as-a-Judge across 3 judges and 14 visual task categories. Our results show that evaluation uncertainty is strongly task-dependent: intervals cover ~40% of the score range for aesthetics and natural images but expand to ~70% for chart and mathematical reasoning, yielding a quantitative reliability map for multimodal evaluation. We further identify a failure mode not captured by standard evaluation metrics, ranking-scoring decoupling, where judges achieve…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

divake/VLM-Judge-Uncertainty
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.