On the Convergence of Policy in Unregularized Policy Mirror Descent

Dachao Lin; Zhihua Zhang

arXiv:2205.08176·math.OC·June 4, 2024

On the Convergence of Policy in Unregularized Policy Mirror Descent

Dachao Lin, Zhihua Zhang

PDF

Open Access

TL;DR

This paper analyzes the convergence behavior of policy in unregularized policy mirror descent using generalized Bregman divergence, revealing conditions for finite-step convergence to optimal policies.

Contribution

It extends previous work by providing convergence rates for policies under generalized Bregman divergence, including classical Euclidean distance, in unregularized policy mirror descent.

Findings

01

Finite-step convergence to optimal policy with certain Bregman divergences

02

Convergence rates established for policies under generalized Bregman divergence

03

Extension of previous convergence results in policy mirror descent

Abstract

In this short note, we give the convergence analysis of the policy in the recent famous policy mirror descent (PMD). We mainly consider the unregularized setting following [11] with generalized Bregman divergence. The difference is that we directly give the convergence rates of policy under generalized Bregman divergence. Our results are inspired by the convergence of value function in previous works and are an extension study of policy mirror descent. Though some results have already appeared in previous work, we further discover a large body of Bregman divergences could give finite-step convergence to an optimal policy, such as the classical Euclidean distance.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStatistical Mechanics and Entropy · Risk and Portfolio Optimization · Nuclear reactor physics and engineering