AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination   Evaluation

Junyang Wang; Yuhang Wang; Guohai Xu; Jing Zhang; Yukai Gu; Haitao; Jia; Jiaqi Wang; Haiyang Xu; Ming Yan; Ji Zhang; Jitao Sang

arXiv:2311.07397·cs.CL·February 26, 2024·5 cites

AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Junyang Wang, Yuhang Wang, Guohai Xu, Jing Zhang, Yukai Gu, Haitao, Jia, Jiaqi Wang, Haiyang Xu, Ming Yan, Ji Zhang, Jitao Sang

PDF

Open Access 1 Repo

TL;DR

AMBER is a novel, low-cost, multi-dimensional benchmark designed to evaluate hallucinations in Multi-modal Large Language Models without relying on large language models, covering various hallucination types and tasks.

Contribution

The paper introduces AMBER, an LLM-free, multi-dimensional benchmark for evaluating hallucinations in MLLMs, addressing high evaluation costs and limited dimensions of previous methods.

Findings

01

Mainstream MLLMs exhibit significant hallucinations.

02

AMBER effectively evaluates hallucinations across multiple dimensions.

03

Guidelines are provided for reducing hallucinations in MLLMs.

Abstract

Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of tasks and hallucinations). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including existence, attribute and relation hallucination. Based on AMBER, we design a low-cost and efficient evaluation pipeline. Additionally, we conduct a comprehensive evaluation and detailed analysis of mainstream MLLMs…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

junyangwang0410/amber
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Text Readability and Simplification