VO$Q$L: Towards Optimal Regret in Model-free RL with Nonlinear Function   Approximation

Alekh Agarwal; Yujia Jin; Tong Zhang

arXiv:2212.06069·cs.LG·December 13, 2022

VO$Q$L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Alekh Agarwal, Yujia Jin, Tong Zhang

PDF

Open Access

TL;DR

This paper introduces VO$Q$L, a new algorithm for model-free reinforcement learning with nonlinear function approximation, achieving near-optimal regret bounds and computational efficiency in linear MDPs.

Contribution

The paper proposes VO$Q$L, a novel, computationally efficient algorithm with provably optimal regret bounds for RL with nonlinear function approximation.

Findings

01

Achieves $ ilde{O}(d ext{sqrt}(HT)+d^6H^{5})$ regret in linear MDPs.

02

First computationally tractable and statistically optimal approach for linear MDPs.

03

Incorporates weighted regression bounds to improve regret performance.

Abstract

We study time-inhomogeneous episodic reinforcement learning (RL) under general function approximation and sparse rewards. We design a new algorithm, Variance-weighted Optimistic $Q$ -Learning (VO $Q$ L), based on $Q$ -learning and bound its regret assuming completeness and bounded Eluder dimension for the regression function class. As a special case, VO $Q$ L achieves $\tilde{O} (d H T + d^{6} H^{5})$ regret over $T$ episodes for a horizon $H$ MDP under ( $d$ -dimensional) linear function approximation, which is asymptotically optimal. Our algorithm incorporates weighted regression-based upper and lower bounds on the optimal value function to obtain this improved regret. The algorithm is computationally efficient given a regression oracle over the function class, making this the first computationally tractable and statistically optimal approach for linear MDPs.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Adversarial Robustness in Machine Learning · Model Reduction and Neural Networks