An operator view of policy gradient methods

Dibya Ghosh; Marlos C. Machado; Nicolas Le Roux

arXiv:2006.11266·cs.LG·October 26, 2020·5 cites

An operator view of policy gradient methods

Dibya Ghosh, Marlos C. Machado, Nicolas Le Roux

PDF

Open Access 1 Video

TL;DR

This paper introduces an operator-based framework for policy gradient methods, providing new insights into their mechanisms, proposing a global lower bound of expected return, and bridging the gap between policy-based and value-based reinforcement learning approaches.

Contribution

It presents a novel operator perspective on policy gradient methods, leading to a better understanding and a new global lower bound for expected return.

Findings

01

Operator framework clarifies policy gradient methods

02

Introduces a new global lower bound of expected return

03

Bridges policy-based and value-based methods

Abstract

We cast policy gradient methods as the repeated application of two operators: a policy improvement operator $I$ , which maps any policy $π$ to a better one $I π$ , and a projection operator $P$ , which finds the best approximation of $I π$ in the set of realizable policies. We use this framework to introduce operator-based versions of traditional policy gradient methods such as REINFORCE and PPO, which leads to a better understanding of their original counterparts. We also use the understanding we develop of the role of $I$ and $P$ to propose a new global lower bound of the expected return. This new perspective allows us to further bridge the gap between policy-based and value-based methods, showing how REINFORCE and the Bellman optimality operator, for example, can be seen as two sides of the same coin.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

An operator view of policy gradient methods· slideslive

Taxonomy

TopicsReinforcement Learning in Robotics · Adaptive Dynamic Programming Control

MethodsEntropy Regularization · Proximal Policy Optimization · REINFORCE