Online Learning in Unknown Markov Games

Yi Tian; Yuanhao Wang; Tiancheng Yu; Suvrit Sra

arXiv:2010.15020·cs.LG·February 9, 2021

Online Learning in Unknown Markov Games

Yi Tian, Yuanhao Wang, Tiancheng Yu, Suvrit Sra

PDF

Open Access 1 Video

TL;DR

This paper introduces the first sublinear regret algorithm for online learning in unknown Markov games, addressing the challenge of unobservable opponent actions and improving regret bounds over previous methods.

Contribution

It proposes a novel algorithm achieving sublinear regret in unknown Markov games, independent of opponents' action space size, and extends analysis to fully observable opponent actions.

Findings

01

Achieves (7K^{2/3}) regret after K episodes.

02

First sublinear regret bound for unknown Markov games.

03

Regret bound is independent of opponents' action space size.

Abstract

We study online learning in unknown Markov games, a problem that arises in episodic multi-agent reinforcement learning where the actions of the opponents are unobservable. We show that in this challenging setting, achieving sublinear regret against the best response in hindsight is statistically hard. We then consider a weaker notion of regret by competing with the \emph{minimax value} of the game, and present an algorithm that achieves a sublinear $\tilde{O} (K^{2/3})$ regret after $K$ episodes. This is the first sublinear regret bound (to our knowledge) for online learning in unknown Markov games. Importantly, our regret bound is independent of the size of the opponents' action spaces. As a result, even when the opponents' actions are fully observable, our regret bound improves upon existing analysis (e.g., (Xie et al., 2020)) by an exponential factor in the number of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Online Learning in Unknown Markov Games· slideslive

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Reinforcement Learning in Robotics · Optimization and Search Problems