Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks

Mahdi Erfanian; Nelson Daniel Troncoso; Aashna Garg; Amabel Gale; Xiaoyu Liu; Pareesa Ameneh Golnari; Shengyu Fu

arXiv:2605.07024·cs.LG·May 12, 2026

Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks

Mahdi Erfanian, Nelson Daniel Troncoso, Aashna Garg, Amabel Gale, Xiaoyu Liu, Pareesa Ameneh Golnari, Shengyu Fu

PDF

1 Repo 1 Datasets

TL;DR

Delulu is a comprehensive multi-lingual benchmark designed to evaluate and detect code hallucinations in fill-in-the-middle tasks of large language models, combining automated and human verification methods.

Contribution

It introduces a verified, adversarially curated benchmark with a novel evaluation pipeline for assessing hallucination detection in code generation models across multiple languages.

Findings

01

The strongest model achieves only 84.5% pass@1 on the benchmark.

02

No model family exceeds 77% similarity in hallucination detection.

03

All tested models produce hallucinations on a significant portion of samples.

Abstract

Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions such as invented API methods, invalid parameters, undefined variables, or non-existent imports. These failures pass superficial review yet introduce runtime errors. We introduce Delulu, a verified multi-lingual benchmark of 1,951 FIM samples across 7 languages and 4 hallucination types. Samples are curated through an adversarial pipeline: a frontier LLM generates plausible hallucinations, four diverse judge models evaluate them, embedding-based clustering mines progressively harder examples, self-contained Docker containers verify that golden completions compile while hallucinated variants produce the expected runtime error, and a final human-expert review removes any remaining biased or trivially decidable samples. We evaluate 11…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

microsoft/delulu
github

Datasets

microsoft/delulu-fim-benchmark
dataset· 749 dl
749 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.