No-regret learning for repeated non-cooperative games with lossy bandits

Wenting Liu; Jinlong Lei; Peng Yi; Yiguang Hong

arXiv:2205.06968·cs.LG·May 17, 2022

No-regret learning for repeated non-cooperative games with lossy bandits

Wenting Liu, Jinlong Lei, Peng Yi, Yiguang Hong

PDF

Open Access

TL;DR

This paper introduces a novel no-regret learning algorithm for repeated non-cooperative games with lossy bandit feedback, demonstrating convergence to Nash equilibrium and applying it to fog computing resource management.

Contribution

It proposes the OGD-lb algorithm for asynchronous online learning in lossy environments, with theoretical convergence guarantees and practical application in fog computing.

Findings

01

Algorithm converges to Nash equilibrium with probability 1.

02

Mean square convergence rate is $ ext{O}(k^{-2eta})$ for strongly monotone games.

03

Numerical experiments validate the algorithm's effectiveness in resource management.

Abstract

This paper considers no-regret learning for repeated continuous-kernel games with lossy bandit feedback. Since it is difficult to give the explicit model of the utility functions in dynamic environments, the players' action can only be learned with bandit feedback. Moreover, because of unreliable communication channels or privacy protection, the bandit feedback may be lost or dropped at random. Therefore, we study the asynchronous online learning strategy of the players to adaptively adjust the next actions for minimizing the long-term regret loss. The paper provides a novel no-regret learning algorithm, called Online Gradient Descent with lossy bandits (OGD-lb). We first give the regret analysis for concave games with differentiable and Lipschitz utilities. Then we show that the action profile converges to a Nash equilibrium with probability 1 when the game is also strictly monotone.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Smart Grid Energy Management · Age of Information Optimization