FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

Betty Xiong; Jillian Fisher; Benjamin Newman; Meng Hu; Shivangi Gupta; Yejin Choi; Lanyan Fang; Russ B Altman

arXiv:2603.19539·cs.CL·March 23, 2026

FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

Betty Xiong, Jillian Fisher, Benjamin Newman, Meng Hu, Shivangi Gupta, Yejin Choi, Lanyan Fang, Russ B Altman

PDF

Open Access

TL;DR

FDARxBench is a comprehensive benchmark for evaluating language models' ability to understand and reason over FDA drug labels, highlighting gaps in factual accuracy, retrieval, and safe refusal.

Contribution

The paper introduces FDARxBench, a novel expert-curated benchmark with a multi-stage QA pipeline for regulatory and clinical reasoning on FDA drug labels.

Findings

01

Models show significant gaps in factual grounding.

02

Long-context retrieval remains challenging.

03

Safe refusal behavior is often inadequate.

Abstract

We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain rich but heterogeneous clinical and regulatory information, making accurate question answering difficult for current language models. In collaboration with FDA regulatory assessors, we introduce FDARxBench, and construct a multi-stage pipeline for generating high-quality, expert curated, QA examples spanning factual, multi-hop, and refusal tasks, and design evaluation protocols to assess both open-book and closed-book reasoning. Experiments across proprietary and open-weight models reveal substantial gaps in factual grounding, long-context retrieval, and safe refusal behavior. While motivated by FDA generic drug assessment needs, this benchmark also…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Biomedical Text Mining and Ontologies · Computational Drug Discovery Methods