Understanding the Pathologies of Approximate Policy Evaluation when   Combined with Greedification in Reinforcement Learning

Kenny Young; Richard S. Sutton

arXiv:2010.15268·cs.LG·October 30, 2020

Understanding the Pathologies of Approximate Policy Evaluation when Combined with Greedification in Reinforcement Learning

Kenny Young, Richard S. Sutton

PDF

Open Access

TL;DR

This paper investigates the pathological behaviors in reinforcement learning algorithms that combine approximate policy evaluation and greedification, revealing how these can lead to convergence to poor policies and policy oscillations, thus challenging existing theoretical guarantees.

Contribution

The paper provides simple examples and analysis showing that RL algorithms with value approximation and greedification can converge to worst policies or oscillate, highlighting their unreliability.

Findings

01

Pathological behaviors like policy oscillation and convergence to worst policies are demonstrated.

02

These behaviors can occur with various function approximations, including neural networks.

03

The results suggest limitations on theoretical guarantees for RL algorithms with value approximation.

Abstract

Despite empirical success, the theory of reinforcement learning (RL) with value function approximation remains fundamentally incomplete. Prior work has identified a variety of pathological behaviours that arise in RL algorithms that combine approximate on-policy evaluation and greedification. One prominent example is policy oscillation, wherein an algorithm may cycle indefinitely between policies, rather than converging to a fixed point. What is not well understood however is the quality of the policies in the region of oscillation. In this paper we present simple examples illustrating that in addition to policy oscillation and multiple fixed points -- the same basic issue can lead to convergence to the worst possible policy for a given approximation. Such behaviours can arise when algorithms optimize evaluation accuracy weighted by the distribution of states that occur under the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Evolutionary Algorithms and Applications · Supply Chain and Inventory Management