Off-Policy Risk Assessment in Markov Decision Processes

Audrey Huang; Liu Leqi; Zachary Chase Lipton; Kamyar Azizzadenesheli

arXiv:2209.10444·cs.LG·September 22, 2022

Off-Policy Risk Assessment in Markov Decision Processes

Audrey Huang, Liu Leqi, Zachary Chase Lipton, Kamyar Azizzadenesheli

PDF

Open Access

TL;DR

This paper extends off-policy risk assessment methods from contextual bandits to Markov decision processes, introducing a doubly robust estimator that reduces variance and improves accuracy in estimating return distributions.

Contribution

It develops the first doubly robust CDF estimator for MDPs, providing theoretical guarantees and practical improvements over importance sampling methods.

Findings

01

Doubly robust estimator significantly reduces variance.

02

Estimator achieves Cramer-Rao lower bound with well-specified models.

03

Experimental results confirm high precision of the proposed method.

Abstract

Addressing such diverse ends as safety alignment with human preferences, and the efficiency of learning, a growing line of reinforcement learning research focuses on risk functionals that depend on the entire distribution of returns. Recent work on \emph{off-policy risk assessment} (OPRA) for contextual bandits introduced consistent estimators for the target policy's CDF of returns along with finite sample guarantees that extend to (and hold simultaneously over) all risk. In this paper, we lift OPRA to Markov decision processes (MDPs), where importance sampling (IS) CDF estimators suffer high variance on longer trajectories due to small effective sample size. To mitigate these problems, we incorporate model-based estimation to develop the first doubly robust (DR) estimator for the CDF of returns in MDPs. This estimator enjoys significantly less variance and, when the model is well…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAge of Information Optimization