Is the Best Better? Bayesian Statistical Model Comparison for Natural   Language Processing

Piotr Szyma\'nski; Kyle Gorman

arXiv:2010.03088·cs.CL·October 8, 2020

Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing

Piotr Szyma\'nski, Kyle Gorman

PDF

TL;DR

This paper introduces a Bayesian model comparison method using k-fold cross-validation to reliably evaluate and rank NLP models across multiple datasets and metrics, addressing concerns about standard split-based evaluations.

Contribution

It presents a novel Bayesian statistical approach for model comparison in NLP, improving upon traditional evaluation methods by considering model uncertainty and multiple datasets.

Findings

01

The method effectively ranks six POS taggers across datasets.

02

It estimates the probability of one model outperforming another.

03

It assesses model equivalence with statistical confidence.

Abstract

Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to estimate the likelihood that one model will outperform the other, or that the two will produce practically equivalent results. We use this technique to rank six English part-of-speech taggers across two data sets and three evaluation metrics.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.