Bayesian Inverse Transition Learning for Offline Settings

Leo Benac; Sonali Parbhoo; Finale Doshi-Velez

arXiv:2308.05075·cs.LG·August 10, 2023

Bayesian Inverse Transition Learning for Offline Settings

Leo Benac, Sonali Parbhoo, Finale Doshi-Velez

PDF

Open Access

TL;DR

This paper introduces a Bayesian inverse transition learning method for offline reinforcement learning, focusing on reliable transition estimation, safety, and uncertainty quantification to improve policy performance and safety.

Contribution

It proposes a new constraint-based Bayesian approach to estimate transition dynamics, reducing variance and enhancing safety and informativeness of policies in offline RL.

Findings

01

High-performing policies with reduced variance

02

Effective uncertainty estimation for safer actions

03

Partial action ranking for improved planning

Abstract

Offline Reinforcement learning is commonly used for sequential decision-making in domains such as healthcare and education, where the rewards are known and the transition dynamics $T$ must be estimated on the basis of batch data. A key challenge for all tasks is how to learn a reliable estimate of the transition dynamics $T$ that produce near-optimal policies that are safe enough so that they never take actions that are far away from the best action with respect to their value functions and informative enough so that they communicate the uncertainties they have. Using data from an expert, we propose a new constraint-based approach that captures our desiderata for reliably learning a posterior distribution of the transition dynamics $T$ that is free from gradients. Our results demonstrate that by using our constraints, we learn a high-performing policy, while considerably reducing the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReservoir Engineering and Simulation Methods · Forecasting Techniques and Applications · Explainable Artificial Intelligence (XAI)