MF-OML: Online Mean-Field Reinforcement Learning with Occupation Measures for Large Population Games

Anran Hu; Junzi Zhang

arXiv:2405.00282·math.OC·September 4, 2025

MF-OML: Online Mean-Field Reinforcement Learning with Occupation Measures for Large Population Games

Anran Hu, Junzi Zhang

PDF

Open Access

TL;DR

This paper introduces MF-OML, a novel online mean-field reinforcement learning algorithm that efficiently computes approximate Nash equilibria in large population symmetric games, with provable guarantees and regret bounds.

Contribution

MF-OML is the first fully polynomial multi-agent RL algorithm for Nash equilibria in large population games beyond zero-sum and potential games.

Findings

01

Achieves high probability regret bounds under monotonicity conditions.

02

Provides the first tractable globally convergent algorithm for monotone mean-field games.

03

Demonstrates effectiveness in large population symmetric games.

Abstract

Reinforcement learning for multi-agent games has attracted lots of attention recently. However, given the challenge of solving Nash equilibria for large population games, existing works with guaranteed polynomial complexities either focus on variants of zero-sum and potential games, or aim at solving (coarse) correlated equilibria, or require access to simulators, or rely on certain assumptions that are hard to verify. This work proposes MF-OML (Mean-Field Occupation-Measure Learning), an online mean-field reinforcement learning algorithm for computing approximate Nash equilibria of large population sequential symmetric games. MF-OML is the first fully polynomial multi-agent reinforcement learning algorithm for provably solving Nash equilibria (up to mean-field approximation gaps that vanish as the number of players $N$ goes to infinity) beyond variants of zero-sum and potential games.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsInnovation Diffusion and Forecasting

MethodsFocus