Boosting Trust Region Policy Optimization by Normalizing Flows Policy

Yunhao Tang; Shipra Agrawal

arXiv:1809.10326·cs.AI·February 4, 2019·20 cites

Boosting Trust Region Policy Optimization by Normalizing Flows Policy

Yunhao Tang, Shipra Agrawal

PDF

Open Access 1 Repo

TL;DR

This paper introduces a normalizing flows policy to enhance trust region policy optimization, enabling better exploration and avoiding local optima, especially in high-dimensional complex tasks.

Contribution

It presents a novel integration of normalizing flows into trust region policy search, improving exploration and performance in complex, high-dimensional environments.

Findings

01

Normalizing flows policy enables better exploration.

02

Significant performance improvements on high-dimensional tasks.

03

Enhanced avoidance of local optima.

Abstract

We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps avoid bad local optima. Through extensive comparisons, we show that the normalizing flows policy significantly improves upon baseline architectures especially on high-dimensional tasks with complex dynamics.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

robintyh1/onpolicybaselines
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Optimization and Search Problems · Stochastic Gradient Optimization Techniques

MethodsNormalizing Flows