A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

Axel Friedrich Wolter; Tobias Sutter

arXiv:2505.04494·math.OC·April 15, 2026

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

Axel Friedrich Wolter, Tobias Sutter

PDF

TL;DR

This paper introduces PGDA-RL, a primal-dual algorithm for reinforcement learning that leverages two-timescale stochastic approximation, enabling online policy updates with theoretical convergence guarantees under weaker assumptions.

Contribution

The paper proposes PGDA-RL, a novel two-timescale primal-dual RL algorithm that operates asynchronously with online updates, removing the need for a simulator or fixed policy, and provides convergence analysis.

Findings

01

PGDA-RL converges almost surely to the optimal policy and value function.

02

Under ergodicity, PGDA-RL achieves a $ ilde{O}(k^{-2/3})$ convergence rate.

03

The algorithm works with correlated data and off-policy information.

Abstract

We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation. Motivated by the challenge of designing algorithms that leverage off-policy data while maintaining on-policy exploration, we propose PGDA-RL, a novel primal-dual Projected Gradient Descent-Ascent algorithm for solving regularized Markov Decision Processes (MDPs). PGDA-RL integrates experience replay-based gradient estimation with a two-timescale decomposition of the underlying nested optimization problem. The algorithm operates asynchronously, interacts with the environment through a single trajectory of correlated data, and updates its policy online in response to the dual variable associated with the occupancy measure of the underlying MDP. We prove that PGDA-RL converges almost surely to the optimal value function and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.