Robust Value Iteration for Continuous Control Tasks

Michael Lutter; Shie Mannor; Jan Peters; Dieter Fox and; Animesh Garg

arXiv:2105.12189·cs.LG·May 27, 2021

Robust Value Iteration for Continuous Control Tasks

Michael Lutter, Shie Mannor, Jan Peters, Dieter Fox and, Animesh Garg

PDF

1 Repo

TL;DR

This paper introduces Robust Fitted Value Iteration, a method that computes robust policies for continuous control tasks by considering adversarial perturbations, improving transferability from simulation to real systems without discretizing states or actions.

Contribution

The paper presents a novel robust value iteration algorithm that incorporates adversarial perturbations in continuous control, derived in closed-form, and demonstrates improved robustness over standard methods.

Findings

01

Robust value iteration outperforms deep RL in transfer tasks.

02

The method does not require state or action discretization.

03

Applied successfully to Furuta pendulum and cartpole.

Abstract

When transferring a control policy from simulation to a physical system, the policy needs to be robust to variations in the dynamics to perform well. Commonly, the optimal policy overfits to the approximate model and the corresponding state-distribution, often resulting in failure to trasnfer underlying distributional shifts. In this paper, we present Robust Fitted Value Iteration, which uses dynamic programming to compute the optimal value function on the compact state domain and incorporates adversarial perturbations of the system dynamics. The adversarial perturbations encourage a optimal policy that is robust to changes in the dynamics. Utilizing the continuous-time perspective of reinforcement learning, we derive the optimal perturbations for the states, actions, observations and model parameters in closed-form. Notably, the resulting algorithm does not require discretization of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

milutter/value_iteration
pytorch

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.