A Short Survey on Probabilistic Reinforcement Learning

Reazul Hasan Russel

arXiv:1901.07010·cs.LG·January 23, 2019·1 cites

A Short Survey on Probabilistic Reinforcement Learning

Reazul Hasan Russel

PDF

Open Access

TL;DR

This paper surveys methods in reinforcement learning that balance exploration and exploitation, focusing on approaches that provide performance guarantees from fixed data samples, especially relevant in sensitive domains.

Contribution

It provides a concise overview of existing techniques for robust policy computation with fixed samples, highlighting approaches that ensure performance guarantees.

Findings

01

Summarizes key methods for exploration-exploitation trade-offs

02

Highlights techniques for robust policy computation from fixed data

03

Discusses applications in sensitive domains

Abstract

A reinforcement learning agent tries to maximize its cumulative payoff by interacting in an unknown environment. It is important for the agent to explore suboptimal actions as well as to pick actions with highest known rewards. Yet, in sensitive domains, collecting more data with exploration is not always possible, but it is important to find a policy with a certain performance guaranty. In this paper, we present a brief survey of methods available in the literature for balancing exploration-exploitation trade off and computing robust solutions from fixed samples in reinforcement learning.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Reinforcement Learning in Robotics · Data Stream Mining Techniques