Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient   Reinforcement Learning

Seungyul Han; Youngchul Sung

arXiv:1905.02363·cs.LG·May 30, 2019·6 cites

Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning

Seungyul Han, Youngchul Sung

PDF

Open Access 1 Repo

TL;DR

This paper introduces a dimension-wise importance sampling weight clipping method for PPO that reduces bias and enhances sample efficiency in high-dimensional action spaces, outperforming existing algorithms.

Contribution

It proposes a novel dimension-wise IS weight clipping technique for PPO, improving learning efficiency and sample reuse in high-dimensional action tasks.

Findings

01

Outperforms PPO in OpenAI Gym tasks

02

Reduces bias in high-dimensional action spaces

03

Enables effective sample reuse

Abstract

In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it difficult to reuse old samples with large IS weights. In this paper, we consider PPO, a representative on-policy algorithm, and propose its improvement by dimension-wise IS weight clipping which separately clips the IS weight of each action dimension to avoid large bias and adaptively controls the IS weight to bound policy update from the current policy. This new technique enables efficient learning for high action-dimensional tasks and reusing of old samples like in off-policy learning to increase the sample efficiency. Numerical results show that the proposed new algorithm…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

seungyulhan/disc
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Machine Learning and ELM · Advanced Bandit Algorithms Research

MethodsEntropy Regularization · Proximal Policy Optimization