Multi-armed Bandit Problem with Known Trend

Djallel Bouneffouf; Rapha\"el Feraud

arXiv:1508.07091·cs.LG·May 15, 2017

Multi-armed Bandit Problem with Known Trend

Djallel Bouneffouf, Rapha\"el Feraud

PDF

Open Access

TL;DR

This paper introduces a variant of the multi-armed bandit problem where the reward trend is known, proposing an adapted algorithm that leverages this information to improve decision-making in online applications.

Contribution

It presents the A-UCB algorithm tailored for known reward trends, with theoretical regret bounds and experimental validation demonstrating its effectiveness.

Findings

01

A-UCB outperforms standard UCB1 in known trend settings.

02

Theoretical regret bounds are improved over traditional methods.

03

Experimental results confirm the advantages of the proposed approach.

Abstract

We consider a variant of the multi-armed bandit model, which we call multi-armed bandit problem with known trend, where the gambler knows the shape of the reward function of each arm but not its distribution. This new problem is motivated by different online problems like active learning, music and interface recommendation applications, where when an arm is sampled by the model the received reward change according to a known trend. By adapting the standard multi-armed bandit algorithm UCB1 to take advantage of this setting, we propose the new algorithm named A-UCB that assumes a stochastic model. We provide upper bounds of the regret which compare favourably with the ones of UCB1. We also confirm that experimentally with different simulations

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Reinforcement Learning in Robotics · Optimization and Search Problems