A new Gradient TD Algorithm with only One Step-size: Convergence Rate   Analysis using $L$-$\lambda$ Smoothness

Hengshuai Yao

arXiv:2307.15892·cs.LG·September 6, 2023

A new Gradient TD Algorithm with only One Step-size: Convergence Rate Analysis using $L$-$\lambda$ Smoothness

Hengshuai Yao

PDF

Open Access

TL;DR

This paper introduces a new single-step-size Gradient TD algorithm called Impression GTD, which achieves faster convergence rates, including linear convergence under $L$-$\lambda$ smoothness, improving upon existing methods for off-policy learning.

Contribution

The paper proposes a truly single-time-scale GTD algorithm with only one step-size parameter and proves it converges at a rate of $O(1/t)$, with linear convergence under $L$-$\lambda$ smoothness, advancing the theoretical understanding of GTD algorithms.

Findings

01

Impression GTD converges faster than existing GTD algorithms.

02

The new algorithm achieves at least $O(1/t)$ convergence rate.

03

Empirical results show significant speedup in convergence on benchmark problems.

Abstract

Gradient Temporal Difference (GTD) algorithms (Sutton et al., 2008, 2009) are the first $O (d)$ ( $d$ is the number features) algorithms that have convergence guarantees for off-policy learning with linear function approximation. Liu et al. (2015) and Dalal et. al. (2018) proved the convergence rates of GTD, GTD2 and TDC are $O (t^{- α /2})$ for some $α \in (0, 1)$ . This bound is tight (Dalal et al., 2020), and slower than $O (1/ t)$ . GTD algorithms also have two step-size parameters, which are difficult to tune. In literature, there is a "single-time-scale" formulation of GTD. However, this formulation still has two step-size parameters. This paper presents a truly single-time-scale GTD algorithm for minimizing the Norm of Expected td Update (NEU) objective, and it has only one step-size parameter. We prove that the new algorithm, called Impression GTD, converges at least as…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning and Algorithms · Reinforcement Learning in Robotics · Age of Information Optimization