Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum   Linear Quadratic Games

Kaiqing Zhang; Zhuoran Yang; Tamer Ba\c{s}ar

arXiv:1906.00729·cs.LG·February 12, 2021·28 cites

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

Kaiqing Zhang, Zhuoran Yang, Tamer Ba\c{s}ar

PDF

Open Access

TL;DR

This paper proves that policy optimization methods for zero-sum linear quadratic games globally converge to Nash equilibria, despite the nonconvex landscape, and introduces algorithms with guaranteed convergence rates.

Contribution

It is the first work to analyze the optimization landscape of LQ games and establish provable convergence of policy optimization to Nash equilibria.

Findings

01

Stationary points in LQ games are Nash equilibria.

02

Developed three algorithms with global sublinear and local linear convergence.

03

Simulation results confirm the algorithms' convergence properties.

Abstract

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave saddle-point problem in the policy space. Specifically, we show that despite its nonconvexity and nonconcavity, zero-sum LQ games have the property that the stationary point of the objective function with respect to the linear feedback control policies constitutes the NE of the game. Building upon this, we develop three projected nested-gradient methods that are guaranteed to converge to the NE of the game. Moreover, we show that all of these algorithms enjoy both globally sublinear and locally linear convergence rates. Simulation results are also provided to illustrate the satisfactory convergence properties of the algorithms. To the best of our knowledge,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Adaptive Dynamic Programming Control · Advanced Bandit Algorithms Research