VIGIL: Tackling Hallucination Detection in Image Recontextualization

Joanna Wojciechowicz; Maria {\L}ubniewska; Jakub Antczak; Justyna Baczy\'nska; Wojciech Gromski; Wojciech Koz{\l}owski; Maciej Zi\k{e}ba

arXiv:2602.14633·cs.CV·February 17, 2026

VIGIL: Tackling Hallucination Detection in Image Recontextualization

Joanna Wojciechowicz, Maria {\L}ubniewska, Jakub Antczak, Justyna Baczy\'nska, Wojciech Gromski, Wojciech Koz{\l}owski, Maciej Zi\k{e}ba

PDF

Open Access

TL;DR

VIGIL introduces a detailed benchmark and detection framework for hallucinations in multimodal image recontextualization, categorizing errors into five types to improve understanding and evaluation of large multimodal models.

Contribution

It provides the first fine-grained categorization and detection pipeline for hallucinations in multimodal image recontextualization tasks, filling a significant gap in model evaluation.

Findings

01

Effective multi-stage detection pipeline demonstrated

02

Comprehensive categorization of hallucination types

03

Open-source release of dataset and tools

Abstract

We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), the first benchmark dataset and framework providing a fine-grained categorization of hallucinations in the multimodal image recontextualization task for large multimodal models (LMMs). While existing research often treats hallucinations as a uniform issue, our work addresses a significant gap in multimodal evaluation by decomposing these errors into five categories: pasted object hallucinations, background hallucinations, object omission, positional & logical inconsistencies, and physical law violations. To address these complexities, we propose a multi-stage detection pipeline. Our architecture processes recontextualized images through a series of specialized steps targeting object-level fidelity, background consistency, and omission detection, leveraging a coordinated ensemble of open-source models, whose…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Generative Adversarial Networks and Image Synthesis · Explainable Artificial Intelligence (XAI)