Robust Imitation via Mirror Descent Inverse Reinforcement Learning

Dong-Sig Han; Hyunseo Kim; Hyundo Lee; Je-Hwan Ryu; Byoung-Tak Zhang

arXiv:2210.11201·cs.LG·January 6, 2023·1 cites

Robust Imitation via Mirror Descent Inverse Reinforcement Learning

Dong-Sig Han, Hyunseo Kim, Hyundo Lee, Je-Hwan Ryu, Byoung-Tak Zhang

PDF

Open Access 1 Video

TL;DR

This paper introduces a robust inverse reinforcement learning method based on mirror descent, which improves reward estimation stability and outperforms existing adversarial IRL techniques across benchmarks.

Contribution

It proposes a novel IRL approach using mirror descent that enhances robustness to reward uncertainty and provides theoretical guarantees with regret bounds.

Findings

01

Outperforms existing adversarial IRL methods in benchmarks

02

Ensures robust reward minimization with regret bound of O(1/T)

03

Provides theoretical proof of stability and convergence

Abstract

Recently, adversarial imitation learning has shown a scalable reward acquisition method for inverse reinforcement learning (IRL) problems. However, estimated reward signals often become uncertain and fail to train a reliable statistical model since the existing methods tend to solve hard optimization problems directly. Inspired by a first-order optimization method called mirror descent, this paper proposes to predict a sequence of reward functions, which are iterative solutions for a constrained convex problem. IRL solutions derived by mirror descent are tolerant to the uncertainty incurred by target density estimation since the amount of reward learning is regulated with respect to local geometric constraints. We prove that the proposed mirror descent update rule ensures robust minimization of a Bregman divergence in terms of a rigorous regret bound of $O (1/ T)$ for step sizes…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Robust Imitation via Mirror Descent Inverse Reinforcement Learning· slideslive

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Receptor Mechanisms and Signaling · Model Reduction and Neural Networks