Statistical Consequences of Dueling Bandits

Nayan Saxena; Pan Chen; Emmy Liu

arXiv:2111.00870·cs.LG·November 2, 2021

Statistical Consequences of Dueling Bandits

Nayan Saxena, Pan Chen, Emmy Liu

PDF

Open Access

TL;DR

This paper examines the statistical properties of dueling bandit algorithms in adaptive experiments, highlighting their strengths in regret minimization and challenges like inflated error rates, informing their practical use.

Contribution

It provides a comparative analysis of traditional uniform sampling and dueling bandit algorithms, revealing their statistical implications in educational and preference-based settings.

Findings

01

Dueling bandit algorithms perform well at cumulative regret minimization.

02

They can lead to inflated Type-I error rates.

03

Reduced statistical power under certain conditions.

Abstract

Multi-Armed-Bandit frameworks have often been used by researchers to assess educational interventions, however, recent work has shown that it is more beneficial for a student to provide qualitative feedback through preference elicitation between different alternatives, making a dueling bandits framework more appropriate. In this paper, we explore the statistical quality of data under this framework by comparing traditional uniform sampling to a dueling bandit algorithm and find that dueling bandit algorithms perform well at cumulative regret minimisation, but lead to inflated Type-I error rates and reduced power under certain circumstances. Through these results we provide insight into the challenges and opportunities in using dueling bandit algorithms to run adaptive experiments.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Model Reduction and Neural Networks · Data Stream Mining Techniques