Span-Based Optimal Sample Complexity for Average Reward MDPs

Matthew Zurek; Yudong Chen

arXiv:2311.13469·cs.LG·March 21, 2024·1 cites

Span-Based Optimal Sample Complexity for Average Reward MDPs

Matthew Zurek, Yudong Chen

PDF

Open Access

TL;DR

This paper establishes a minimax optimal sample complexity bound for learning near-optimal policies in average-reward MDPs, improving theoretical understanding and reducing dependence on mixing assumptions.

Contribution

It introduces a new reduction from average-reward to discounted MDPs and provides improved bounds for discounted MDPs, achieving optimal sample complexity dependence on key parameters.

Findings

01

Achieves minimax optimal sample complexity bound $ ilde{O}(SAH/\varepsilon^2)$

02

Develops tighter bounds for variance parameters in terms of the span of the bias function

03

Circumvents known lower bounds for discounted MDPs under certain discount regimes

Abstract

We study the sample complexity of learning an $ε$ -optimal policy in an average-reward Markov decision process (MDP) under a generative model. We establish the complexity bound $O (S A \frac{H}{ε ^{2}})$ , where $H$ is the span of the bias function of the optimal policy and $S A$ is the cardinality of the state-action space. Our result is the first that is minimax optimal (up to log factors) in all parameters $S, A, H$ and $ε$ , improving on existing work that either assumes uniformly bounded mixing times for all policies or has suboptimal dependence on the parameters. Our result is based on reducing the average-reward MDP to a discounted MDP. To establish the optimality of this reduction, we develop improved bounds for $γ$ -discounted MDPs, showing that $O (S A \frac{H}{( 1 - γ ) ^{2} ε ^{2}})$ samples suffice to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Machine Learning and Algorithms