An Asymptotically Optimal Batched Algorithm for the Dueling Bandit   Problem

Arpit Agarwal; Rohan Ghuge; Viswanath Nagarajan

arXiv:2209.12108·cs.LG·September 27, 2022

An Asymptotically Optimal Batched Algorithm for the Dueling Bandit Problem

Arpit Agarwal, Rohan Ghuge, Viswanath Nagarajan

PDF

Open Access

TL;DR

This paper introduces a batched algorithm for the dueling bandit problem that achieves asymptotic optimal regret with only logarithmic rounds, making it suitable for large-scale applications where sequential updates are impractical.

Contribution

The paper presents the first asymptotically optimal batched algorithm for the dueling bandit problem under the Condorcet condition, matching sequential regret bounds with significantly fewer rounds.

Findings

01

Achieves asymptotic regret of O(K^2 log^2(K)) + O(K log T) in O(log T) rounds.

02

Performs nearly as well as fully sequential algorithms in real-world datasets.

03

Uses only logarithmic rounds, reducing the need for frequent updates in large-scale systems.

Abstract

We study the $K$ -armed dueling bandit problem, a variation of the traditional multi-armed bandit problem in which feedback is obtained in the form of pairwise comparisons. Previous learning algorithms have focused on the $fully adaptive$ setting, where the algorithm can make updates after every comparison. The "batched" dueling bandit problem is motivated by large-scale applications like web search ranking and recommendation systems, where performing sequential updates may be infeasible. In this work, we ask: $\textit{is there a solution using only a few adaptive rounds that matches the asymptotic regret bounds of the best sequential algorithms for$ K $-armed dueling bandits?}$ We answer this in the affirmative $under the Condorcet condition$ , a standard setting of the $K$ -armed dueling bandit problem. We obtain asymptotic regret of $O (K^{2} lo g^{2} (K)) + O (K lo g (T))$ in…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Optimization and Search Problems · Machine Learning and Algorithms