VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models

Hyesu Lim; Jinho Choi; Taekyung Kim; Byeongho Heo; Jaegul Choo; Dongyoon Han

arXiv:2603.07335·cs.AI·March 10, 2026

VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models

Hyesu Lim, Jinho Choi, Taekyung Kim, Byeongho Heo, Jaegul Choo, Dongyoon Han

PDF

Open Access

TL;DR

VisualScratchpad is an interactive tool that analyzes visual concepts in vision language models during inference, helping to understand and debug model failures by linking visual concepts to text tokens.

Contribution

It introduces VisualScratchpad, a novel interface that enables systematic visual concept analysis and causal debugging in vision language models during inference.

Findings

01

Reveals failure modes like limited cross-modal alignment.

02

Identifies misleading visual concepts affecting model outputs.

03

Provides a tool for systematic debugging of vision language models.

Abstract

High-performing vision language models still produce incorrect answers, yet their failure modes are often difficult to explain. To make model internals more accessible and enable systematic debugging, we introduce VisualScratchpad, an interactive interface for visual concept analysis during inference. We apply sparse autoencoders to the vision encoder and link the resulting visual concepts to text tokens via text-to-image attention, allowing us to examine which visual concepts are both captured by the vision encoder and utilized by the language model. VisualScratchpad also provides a token-latent heatmap view that suggests a sufficient set of latents for effective concept ablation in causal analysis. Through case studies, we reveal three underexplored failure modes: limited cross-modal alignment, misleading visual concepts, and unused hidden cues. Project page:…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Generative Adversarial Networks and Image Synthesis · Explainable Artificial Intelligence (XAI)