Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level   Alignment

Pritam Sarkar; Sayna Ebrahimi; Ali Etemad; Ahmad Beirami; Sercan \"O.; Ar{\i}k; Tomas Pfister

arXiv:2405.18654·cs.CV·March 3, 2025·3 cites

Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Pritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami, Sercan \"O., Ar{\i}k, Tomas Pfister

PDF

Open Access 3 Models

TL;DR

This paper introduces Data-augmented Phrase-level Alignment (DPA), a novel training method that significantly reduces object hallucinations in Multimodal Large Language Models while maintaining their general capabilities.

Contribution

The work proposes DPA, a new loss function that leverages generated phrase-level data augmentation to mitigate hallucinations in MLLMs without sacrificing overall performance.

Findings

01

DPA reduces hallucination rates by up to 4.2% on image description tasks.

02

Finetuned MLLMs with DPA, called HALVA, improve F1 scores by up to 13.4% on hallucination visual question-answering.

03

DPA preserves the general vision-language capabilities of MLLMs.

Abstract

Despite their significant advancements, Multimodal Large Language Models (MLLMs) often generate factually inaccurate information, referred to as hallucination. In this work, we address object hallucinations in MLLMs, where information is generated about an object not present in the input image. We introduce Data-augmented Phrase-level Alignment (DPA), a novel loss which can be applied to instruction-tuned off-the-shelf MLLMs to mitigate hallucinations, while preserving their general vision-language capabilities. To fine-tune MLLMs with DPA, we first generate a set of `hallucinated' and `correct' response pairs through generative data augmentation by selectively altering the ground-truth information of the correct responses at a phrase level. The DPA loss is then used to train MLLMs to reduce the likelihood of hallucinated phrases compared to the correct ones. Our thorough evaluation on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsFunctional Brain Connectivity Studies · EEG and Brain-Computer Interfaces · Neurological disorders and treatments

MethodsSparse Evolutionary Training