The Predictron: End-To-End Learning and Planning

David Silver; Hado van Hasselt; Matteo Hessel; Tom Schaul; Arthur; Guez; Tim Harley; Gabriel Dulac-Arnold; David Reichert; Neil Rabinowitz,; Andre Barreto; Thomas Degris

arXiv:1612.08810·cs.LG·July 21, 2017·89 cites

The Predictron: End-To-End Learning and Planning

David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur, Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz,, Andre Barreto, Thomas Degris

PDF

Open Access 1 Repo

TL;DR

The paper introduces the predictron, an end-to-end trainable architecture that models planning as a Markov reward process, enabling more accurate value predictions in complex environments.

Contribution

It presents the predictron architecture, a novel fully abstract model that performs multi-step imagined planning within a neural network framework.

Findings

01

Outperforms conventional neural networks in maze and pool simulations

02

Accurately approximates true value functions through multi-step internal predictions

03

Demonstrates effectiveness in procedurally generated environments

Abstract

One of the key challenges of artificial intelligence is to learn models that are effective in the context of planning. In this document we introduce the predictron architecture. The predictron consists of a fully abstract model, represented by a Markov reward process, that can be rolled forward multiple "imagined" planning steps. Each forward pass of the predictron accumulates internal rewards and values over multiple planning depths. The predictron is trained end-to-end so as to make these accumulated values accurately approximate the true value function. We applied the predictron to procedurally generated random mazes and a simulator for the game of pool. The predictron yielded significantly more accurate predictions than conventional deep neural network architectures.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

zhongwen/predictron
tf

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · AI-based Problem Solving and Planning · Evolutionary Algorithms and Applications