Multi-agent Natural Actor-critic Reinforcement Learning Algorithms

Prashant Trivedi; Nandyala Hemachandra

arXiv:2109.01654·cs.LG·April 5, 2022

Multi-agent Natural Actor-critic Reinforcement Learning Algorithms

Prashant Trivedi, Nandyala Hemachandra

PDF

Open Access

TL;DR

This paper introduces three decentralized multi-agent natural actor-critic algorithms that optimize joint policies in multi-agent reinforcement learning, with proven convergence and practical effectiveness demonstrated in traffic and multi-agent scenarios.

Contribution

It presents the first convergence proofs for fully decentralized multi-agent natural actor-critic algorithms with linear function approximation.

Findings

01

Algorithms converge to stable policy sets.

02

Achieved 25% reduction in traffic congestion.

03

Performance comparable to existing methods in multi-agent tasks.

Abstract

Multi-agent actor-critic algorithms are an important part of the Reinforcement Learning paradigm. We propose three fully decentralized multi-agent natural actor-critic (MAN) algorithms in this work. The objective is to collectively find a joint policy that maximizes the average long-term return of these agents. In the absence of a central controller and to preserve privacy, agents communicate some information to their neighbors via a time-varying communication network. We prove convergence of all the 3 MAN algorithms to a globally asymptotically stable set of the ODE corresponding to actor update; these use linear function approximations. We show that the Kullback-Leibler divergence between policies of successive iterates is proportional to the objective function's gradient. We observe that the minimum singular value of the Fisher information matrix is well within the reciprocal of the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics