Quantile-Based Deep Reinforcement Learning using Two-Timescale Policy   Gradient Algorithms

Jinyang Jiang; Jiaqiao Hu; and Yijie Peng

arXiv:2305.07248·cs.LG·May 15, 2023·1 cites

Quantile-Based Deep Reinforcement Learning using Two-Timescale Policy Gradient Algorithms

Jinyang Jiang, Jiaqiao Hu, and Yijie Peng

PDF

Open Access 1 Repo

TL;DR

This paper introduces novel deep reinforcement learning algorithms, QPO and QPPO, that optimize the quantile of cumulative rewards using two-timescale policy gradient methods, outperforming existing baselines.

Contribution

The paper proposes the first neural network-based algorithms for quantile optimization in deep RL, utilizing two-timescale updates for improved performance.

Findings

01

QPO and QPPO outperform baseline algorithms under the quantile criterion.

02

QPPO achieves higher efficiency with multiple updates per episode.

03

Algorithms effectively optimize the quantile of cumulative rewards.

Abstract

Classical reinforcement learning (RL) aims to optimize the expected cumulative reward. In this work, we consider the RL setting where the goal is to optimize the quantile of the cumulative reward. We parameterize the policy controlling actions by neural networks, and propose a novel policy gradient algorithm called Quantile-Based Policy Optimization (QPO) and its variant Quantile-Based Proximal Policy Optimization (QPPO) for solving deep RL problems with quantile objectives. QPO uses two coupled iterations running at different timescales for simultaneously updating quantiles and policy parameters, whereas QPPO is an off-policy version of QPO that allows multiple updates of parameters during one simulation episode, leading to improved algorithm efficiency. Our numerical results indicate that the proposed algorithms outperform the existing baseline algorithms under the quantile criterion.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

jinyangjiangai/quantile-based-policy-optimization
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Supply Chain and Inventory Management · Advanced Multi-Objective Optimization Algorithms