Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)

Chenhao Fang; Jordi Mola; Mark Harman; Jason Nawrocki; Vaibhav Shrivastava; Yue Cheng; Jay Minesh Shah; Katayoun Zand; Mansi Tripathi; Arya Pudota; Matthew Becker; Herv\'e Robert; Abhishek Gulati

arXiv:2604.11141·cs.LG·April 14, 2026

Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)

Chenhao Fang, Jordi Mola, Mark Harman, Jason Nawrocki, Vaibhav Shrivastava, Yue Cheng, Jay Minesh Shah, Katayoun Zand, Mansi Tripathi, Arya Pudota, Matthew Becker, Herv\'e Robert, Abhishek Gulati

PDF

TL;DR

This paper introduces HUMBR, a hybrid utility approach to significantly reduce hallucinations in enterprise AI workflows by combining semantic and lexical methods, validated through extensive benchmarks and real-world deployment.

Contribution

We propose a novel Hybrid Utility MBR framework that synthesizes semantic and lexical techniques, providing rigorous error bounds and demonstrating superior performance over existing methods.

Findings

01

81% of suggestions were preferred over human ground truth

02

HUMBR significantly outperforms standard Universal Self-Consistency

03

Critical recall failures were virtually eliminated

Abstract

Although LLMs drive automation, it is critical to ensure immense consideration for high-stakes enterprise workflows such as those involving legal matters, risk management, and privacy compliance. For Meta, and other organizations like ours, a single hallucinated clause in such high stakes workflows risks material consequences. We show that by framing hallucination mitigation as a Minimum Bayes Risk (MBR) problem, we can dramatically reduce this risk. Specifically, we introduce a Hybrid Utility MBR (HUMBR) framework that synthesizes semantic embedding similarity with lexical precision to identify consensus without ground-truth references, for which we derive rigorous error bounds. We complement this theoretical analysis with a comprehensive empirical evaluation on widely-used public benchmark suites (TruthfulQA and LegalBench) and also real world data from Meta production deployment. The…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.