Natural Gradient Deep Q-learning

Ethan Knight; Osher Lerner

arXiv:1803.07482·cs.LG·November 15, 2018·5 cites

Natural Gradient Deep Q-learning

Ethan Knight, Osher Lerner

PDF

Open Access

TL;DR

This paper introduces NGDQN, a natural-gradient-based deep Q-learning algorithm that stabilizes training, reduces hyperparameter sensitivity, and outperforms or matches traditional DQN methods across control tasks.

Contribution

It presents a novel natural-gradient approach for deep Q-learning, demonstrating improved stability and reduced hyperparameter sensitivity without using target networks.

Findings

01

NGDQN outperforms DQN without target networks.

02

NGDQN matches DQN with target networks in performance.

03

NGDQN is less sensitive to hyperparameter tuning.

Abstract

We present a novel algorithm to train a deep Q-learning agent using natural-gradient techniques. We compare the original deep Q-network (DQN) algorithm to its natural-gradient counterpart, which we refer to as NGDQN, on a collection of classic control domains. Without employing target networks, NGDQN significantly outperforms DQN without target networks, and performs no worse than DQN with target networks, suggesting that NGDQN stabilizes training and can help reduce the need for additional hyperparameter tuning. We also find that NGDQN is less sensitive to hyperparameter optimization relative to DQN. Together these results suggest that natural-gradient techniques can improve value-function optimization in deep reinforcement learning.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications · Domain Adaptation and Few-Shot Learning · Reinforcement Learning in Robotics

MethodsDense Connections · Convolution · Q-Learning · Deep Q-Network