A0C: Alpha Zero in Continuous Action Space

Thomas M. Moerland; Joost Broekens; Aske Plaat; Catholijn M. Jonker

arXiv:1805.09613·stat.ML·May 25, 2018·20 cites

A0C: Alpha Zero in Continuous Action Space

Thomas M. Moerland, Joost Broekens, Aske Plaat, Catholijn M. Jonker

PDF

Open Access 2 Repos

TL;DR

This paper extends Alpha Zero's tree search and deep learning framework to handle continuous action spaces, demonstrating preliminary success in robotic control tasks and paving the way for broader real-world applications.

Contribution

It introduces theoretical modifications to Alpha Zero for continuous actions and provides initial experimental validation on a control task.

Findings

01

Feasibility of Alpha Zero extension to continuous actions

02

Preliminary success on Pendulum swing-up task

03

First step towards real-world continuous domain applications

Abstract

A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, navigation and self-driving cars. This paper presents the necessary theoretical extensions of Alpha Zero to deal with continuous action space. We also provide some preliminary experiments on the Pendulum swing-up task, empirically showing the feasibility of our approach. Thereby, this work provides a first step towards the application of iterated search and learning in domains with a continuous action space.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsArtificial Intelligence in Games · Reinforcement Learning in Robotics · Advanced Bandit Algorithms Research