Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

Asaf Cassel; Haipeng Luo; Aviv Rosenberg; Dmitry Sotnikov

arXiv:2405.07637·cs.LG·May 15, 2024

Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

Asaf Cassel, Haipeng Luo, Aviv Rosenberg, Dmitry Sotnikov

PDF

Open Access

TL;DR

This paper advances reinforcement learning in scenarios where only aggregate rewards at episode end are available, extending previous tabular results to linear function approximation with near-optimal regret guarantees.

Contribution

It introduces the first algorithms for RL with aggregate bandit feedback in linear MDPs, achieving near-optimal regret bounds.

Findings

01

Developed a value-based optimistic algorithm with a new randomization technique.

02

Created a policy optimization algorithm using a novel hedging scheme.

03

Achieved near-optimal regret guarantees in linear MDPs with aggregate feedback.

Abstract

In many real-world applications, it is hard to provide a reward signal in each step of a Reinforcement Learning (RL) process and more natural to give feedback when an episode ends. To this end, we study the recently proposed model of RL with Aggregate Bandit Feedback (RL-ABF), where the agent only observes the sum of rewards at the end of an episode instead of each reward individually. Prior work studied RL-ABF only in tabular settings, where the number of states is assumed to be small. In this paper, we extend ABF to linear function approximation and develop two efficient algorithms with near-optimal regret guarantees: a value-based optimistic algorithm built on a new randomization technique with a Q-functions ensemble, and a policy optimization algorithm that uses a novel hedging scheme over the ensemble.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Machine Learning and ELM · Distributed Sensor Networks and Detection Algorithms