Slowly Changing Adversarial Bandit Algorithms are Efficient for   Discounted MDPs

Ian A. Kash; Lev Reyzin; Zishun Yu

arXiv:2205.09056·cs.LG·March 12, 2024

Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs

Ian A. Kash, Lev Reyzin, Zishun Yu

PDF

Open Access

TL;DR

This paper presents a reduction from discounted infinite-horizon reinforcement learning to multi-armed bandits, showing that slowly changing adversarial bandit algorithms can efficiently solve discounted MDPs under certain conditions.

Contribution

It introduces a black-box reduction framework that leverages adversarial bandit algorithms for efficient reinforcement learning in discounted MDPs.

Findings

01

Adversarial bandit algorithms achieve optimal regret in discounted MDPs.

02

The reduction works under ergodicity and fast mixing assumptions.

03

Exponential-weight algorithm performs well in this setting.

Abstract

Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement learning to multi-armed bandits, where, specifically, an independent bandit learner is placed in each state. We show that, under ergodicity and fast mixing assumptions, any slowly changing adversarial bandit algorithm achieving optimal regret in the adversarial bandit setting can also attain optimal expected regret in infinite-horizon discounted Markov decision processes, with respect to the number of rounds $T$ . Furthermore, we examine our reduction using a specific instance of the exponential-weight algorithm.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Reinforcement Learning in Robotics · Advanced Bandit Algorithms Research