Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

Marcel Gr\"opl; Jaewoo Jung; Seungryong Kim; Marc Pollefeys; Sunghwan Hong

arXiv:2604.08456·cs.CV·April 10, 2026

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

Marcel Gr\"opl, Jaewoo Jung, Seungryong Kim, Marc Pollefeys, Sunghwan Hong

PDF

TL;DR

This paper introduces a training-free, entropy-gradient based grounding method for vision-language models that improves evidence retrieval and interpretability, especially for detail-critical and high-resolution tasks.

Contribution

It proposes a novel, training-free approach using entropy gradients for evidence retrieval and multi-region support, enhancing interpretability and performance in vision-language models.

Findings

01

Consistent improvements across seven benchmarks and four architectures.

02

Largest gains observed on detail-critical and high-resolution tasks.

03

Produces more interpretable evidence localizations.

Abstract

Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread across multiple regions, as in documents and compositional queries. We address this by framing grounding as test-time evidence retrieval: given a query, the model should actively identify where to look next to resolve ambiguity. To this end, we propose a training-free, model-intrinsic grounding method that uses uncertainty as supervision. Specifically, we compute the entropy of the model's next-token distribution and backpropagate it to the visual token embeddings to obtain an entropy-gradient relevance map, without auxiliary detectors or attention-map heuristics. We then extract and rank multiple coherent regions to support multi-evidence queries, and introduce an iterative zoom-and-reground procedure with a spatial-entropy stopping rule to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.