Rethinking the Global Convergence of Softmax Policy Gradient with Linear   Function Approximation

Max Qiushi Lin; Jincheng Mei; Matin Aghaei; Michael Lu; Bo Dai; Alekh; Agarwal; Dale Schuurmans; Csaba Szepesvari; Sharan Vaswani

arXiv:2505.03155·cs.LG·May 7, 2025

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

Max Qiushi Lin, Jincheng Mei, Matin Aghaei, Michael Lu, Bo Dai, Alekh, Agarwal, Dale Schuurmans, Csaba Szepesvari, Sharan Vaswani

PDF

Open Access

TL;DR

This paper reexamines the convergence properties of Softmax Policy Gradient methods with linear function approximation, showing that approximation error does not hinder global convergence and establishing conditions for guaranteed optimality.

Contribution

It demonstrates that approximation error is irrelevant for convergence in Lin-SPG and identifies feature conditions ensuring asymptotic global convergence.

Findings

01

Approximation error does not affect convergence in Lin-SPG.

02

Under certain feature conditions, Lin-SPG converges to the optimal policy.

03

Constant learning rates can also ensure convergence.

Abstract

Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typically used with function approximation. In this setting, the approximation error in modeling problem-dependent quantities is a key notion for characterizing the global convergence of PG methods. We focus on Softmax PG with linear function approximation (referred to as $Lin-SPG$ ) and demonstrate that the approximation error is irrelevant to the algorithm's global convergence even for the stochastic bandit setting. Consequently, we first identify the necessary and sufficient conditions on the feature representation that can guarantee the asymptotic global convergence of $Lin-SPG$ . Under these feature conditions, we prove that $T$ iterations of $Lin-SPG$ with a problem-specific learning…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic processes and financial applications · Economic theories and models · Stochastic Gradient Optimization Techniques

MethodsFocus · Softmax