Improved Regret for Efficient Online Reinforcement Learning with Linear   Function Approximation

Uri Sherman; Tomer Koren; Yishay Mansour

arXiv:2301.13087·cs.LG·January 31, 2023·1 cites

Improved Regret for Efficient Online Reinforcement Learning with Linear Function Approximation

Uri Sherman, Tomer Koren, Yishay Mansour

PDF

Open Access 1 Video

TL;DR

This paper introduces a new efficient algorithm for online reinforcement learning with linear function approximation under adversarial costs, achieving significantly improved regret bounds in challenging settings.

Contribution

It presents a novel policy optimization algorithm combining mirror descent and least squares evaluation, with improved regret bounds for unknown dynamics and bandit feedback.

Findings

01

Achieves $ ilde{O}(K^{6/7})$ regret without simulator.

02

Achieves $ ilde{O}(K^{2/3})$ regret with a simulator.

03

Outperforms previous state-of-the-art regret bounds.

Abstract

We study reinforcement learning with linear function approximation and adversarially changing cost functions, a setup that has mostly been considered under simplifying assumptions such as full information feedback or exploratory conditions.We present a computationally efficient policy optimization algorithm for the challenging general setting of unknown dynamics and bandit feedback, featuring a combination of mirror-descent and least squares policy evaluation in an auxiliary MDP used to compute exploration bonuses.Our algorithm obtains an $O (K^{6/7})$ regret bound, improving significantly over previous state-of-the-art of $O (K^{14/15})$ in this setting. In addition, we present a version of the same algorithm under the assumption a simulator of the environment is available to the learner (but otherwise no exploratory assumptions are made), and prove it obtains…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Improved Regret for Efficient Online Reinforcement Learning with Linear Function Approximation· slideslive

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Adversarial Robustness in Machine Learning · Reinforcement Learning in Robotics