Unraveling overoptimism and publication bias in ML-driven science

Pouria Saidi; Gautam Dasarathy; Visar Berisha

arXiv:2405.14422·cs.LG·July 15, 2024·1 cites

Unraveling overoptimism and publication bias in ML-driven science

Pouria Saidi, Gautam Dasarathy, Visar Berisha

PDF

Open Access

TL;DR

This paper examines overoptimism and publication bias in ML research, introducing a stochastic model to correct biases and provide realistic performance estimates, with applications to neurological classification studies.

Contribution

It presents a novel stochastic model that accounts for overfitting and publication bias, enabling more accurate estimation of true ML performance from published results.

Findings

01

The model effectively estimates underlying learning curves.

02

Corrected performance estimates are more realistic.

03

Application reveals inherent limits of ML in neurological classification.

Abstract

Machine Learning (ML) is increasingly used across many disciplines with impressive reported results. However, recent studies suggest published performance of ML models are often overoptimistic. Validity concerns are underscored by findings of an inverse relationship between sample size and reported accuracy in published ML models, contrasting with the theory of learning curves where accuracy should improve or remain stable with increasing sample size. This paper investigates factors contributing to overoptimism in ML-driven science, focusing on overfitting and publication bias. We introduce a novel stochastic model for observed accuracy, integrating parametric learning curves and the aforementioned biases. We construct an estimator that corrects for these biases in observed data. Theoretical and empirical results show that our framework can estimate the underlying learning curve,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsScientific Computing and Data Management