Greedy-based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning

Lipeng Wan; Zeyang Liu; Xingyu Chen; Han Wang; Xuguang Lan

arXiv:2112.04454·cs.MA·March 5, 2026

Greedy-based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning

Lipeng Wan, Zeyang Liu, Xingyu Chen, Han Wang, Xuguang Lan

PDF

Open Access

TL;DR

This paper introduces GVR, a novel value representation method for multi-agent reinforcement learning that ensures optimal consistency by transforming the joint Q value function, outperforming existing methods in benchmarks.

Contribution

The paper derives the joint Q value function expression for LVD and MVD, and proposes GVR to ensure optimal consistency through inferior target shaping and experience replay.

Findings

01

GVR guarantees optimal consistency with sufficient exploration.

02

GVR outperforms state-of-the-art baselines on various benchmarks.

03

Theoretical proofs confirm GVR's effectiveness in matrix games.

Abstract

Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the maximal true Q value). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and further eliminates the non-optimal STNs via superior experience replay. In addition, GVR…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics