BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

Anna Sokol; Elizabeth Daly; Michael Hind; David Piorkowski; Xiangliang Zhang; Nuno Moniz; Nitesh Chawla

arXiv:2410.12974·cs.CL·June 4, 2025

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

Anna Sokol, Elizabeth Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, Nitesh Chawla

PDF

Open Access 2 Datasets 1 Video

TL;DR

BenchmarkCards provides a standardized documentation framework for LLM benchmarks, improving transparency, comparability, and ease of selection for users by systematically capturing key benchmark attributes.

Contribution

It introduces a validated, standardized documentation framework for LLM benchmarks, addressing complexity and transparency issues in benchmark selection.

Findings

01

Simplifies benchmark selection process.

02

Enhances transparency and comparability.

03

Validated through user studies.

Abstract

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult given the many available options. This complexity not only increases the risk of benchmark misuse and misinterpretation but also demands substantial effort from LLM users, seeking the most suitable benchmarks for their specific needs. To address these issues, we introduce \texttt{BenchmarkCards}, an intuitive and validated documentation framework that standardizes critical benchmark attributes such as objectives, methodologies, data sources, and limitations. Through user studies involving benchmark creators and users, we show that \texttt{BenchmarkCards} can simplify benchmark selection and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Datasets

Videos

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks· slideslive

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling