Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

Guangzhi Xiong; Eric Xie; Corey Williams; Myles Kim; Amir Hassan Shariatmadari; Sikun Guo; Stefan Bekiranov; Aidong Zhang

arXiv:2505.14599·cs.CL·June 10, 2025

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

Guangzhi Xiong, Eric Xie, Corey Williams, Myles Kim, Amir Hassan Shariatmadari, Sikun Guo, Stefan Bekiranov, Aidong Zhang

PDF

Open Access 1 Repo 3 Datasets

TL;DR

This paper introduces benchmarks and tools to evaluate and improve the truthfulness and factual grounding of hypotheses generated by large language models in scientific research, addressing hallucination issues.

Contribution

It presents TruthHypo, a benchmark for hypothesis truthfulness, and KnowHD, a hallucination detector, advancing systematic evaluation of LLMs in scientific hypothesis generation.

Findings

01

LLMs struggle to generate truthful hypotheses.

02

KnowHD effectively filters truthful hypotheses.

03

Human evaluation confirms the utility of KnowHD.

Abstract

Large language models (LLMs) have shown significant potential in scientific disciplines such as biomedicine, particularly in hypothesis generation, where they can analyze vast literature, identify patterns, and suggest research directions. However, a key challenge lies in evaluating the truthfulness of generated hypotheses, as verifying their accuracy often requires substantial time and resources. Additionally, the hallucination problem in LLMs can lead to the generation of hypotheses that appear plausible but are ultimately incorrect, undermining their reliability. To facilitate the systematic study of these challenges, we introduce TruthHypo, a benchmark for assessing the capabilities of LLMs in generating truthful scientific hypotheses, and KnowHD, a knowledge-based hallucination detector to evaluate how well hypotheses are grounded in existing knowledge. Our results show that LLMs…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

teddy-xionggz/truthhypo
pytorchOfficial

Datasets

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Artificial Intelligence in Healthcare and Education · Explainable Artificial Intelligence (XAI)