Predictor-Corrector(PC) Temporal Difference(TD) Learning (PCTD)

Caleb Bowyer

arXiv:2104.09620·cs.LG·April 21, 2021·1 cites

Predictor-Corrector(PC) Temporal Difference(TD) Learning (PCTD)

Caleb Bowyer

PDF

Open Access

TL;DR

This paper introduces Predictor-Corrector Temporal Difference (PCTD), a novel RL algorithm inspired by numerical ODE methods, which improves approximation accuracy and reduces error in value function estimation compared to traditional TD methods.

Contribution

It proposes a new class of TD learning algorithms based on predictor-corrector methods from numerical analysis, providing both causal and non-causal implementations with improved error bounds.

Findings

01

PCTD algorithms show reduced Taylor Series error in parameter approximation.

02

Simulation results demonstrate PCTD's superior performance over TD(0) in infinite horizon tasks.

03

Both causal and non-causal PCTD implementations are viable and effective.

Abstract

Using insight from numerical approximation of ODEs and the problem formulation and solution methodology of TD learning through a Galerkin relaxation, I propose a new class of TD learning algorithms. After applying the improved numerical methods, the parameter being approximated has a guaranteed order of magnitude reduction in the Taylor Series error of the solution to the ODE for the parameter $θ (t)$ that is used in constructing the linearly parameterized value function. Predictor-Corrector Temporal Difference (PCTD) is what I call the translated discrete time Reinforcement Learning(RL) algorithm from the continuous time ODE using the theory of Stochastic Approximation(SA). Both causal and non-causal implementations of the algorithm are provided, and simulation results are listed for an infinite horizon task to compare the original TD(0) algorithm against both versions of PCTD(0).

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Neural Networks and Applications · Model Reduction and Neural Networks