Inverse reinforcement learning by expert imitation for the stochastic   linear-quadratic optimal control problem

Zhongshi Sun; Guangyan Jia

arXiv:2405.17085·math.OC·May 28, 2024·Neurocomputing

Inverse reinforcement learning by expert imitation for the stochastic linear-quadratic optimal control problem

Zhongshi Sun, Guangyan Jia

PDF

Open Access

TL;DR

This paper introduces model-based and model-free inverse reinforcement learning algorithms for stochastic linear-quadratic control, enabling imitation of expert behavior without prior knowledge of the system, with proven convergence and verified effectiveness.

Contribution

It presents novel IRL algorithms for stochastic LQ control, including a model-free approach that requires only behavior data, along with convergence and stability proofs.

Findings

01

The algorithms successfully imitate expert policies.

02

The model-free IRL performs well without system identification.

03

Convergence and stability are theoretically guaranteed.

Abstract

This article studies inverse reinforcement learning (IRL) for the stochastic linear-quadratic optimal control problem, where two agents are considered. A learner agent does not know the expert agent's performance cost function, but it imitates the behavior of the expert agent by constructing an underlying cost function that obtains the same optimal feedback control as the expert's. We first develop a model-based IRL algorithm, which consists of a policy correction and a policy update from the policy iteration in reinforcement learning, as well as a cost function weight reconstruction based on the inverse optimal control. Then, under this scheme, we propose a model-free off-policy IRL algorithm, which does not need to know or identify the system and only needs to collect the behavior data of the expert agent and learner agent once during the iteration process. Moreover, the proofs of the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdaptive Dynamic Programming Control