Learning Risk-Aware Quadrupedal Locomotion using Distributional   Reinforcement Learning

Lukas Schneider; Jonas Frey; Takahiro Miki; Marco Hutter

arXiv:2309.14246·cs.RO·May 6, 2024

Learning Risk-Aware Quadrupedal Locomotion using Distributional Reinforcement Learning

Lukas Schneider, Jonas Frey, Takahiro Miki, Marco Hutter

PDF

TL;DR

This paper introduces a novel risk-aware reinforcement learning method for quadrupedal robots that explicitly models and adjusts risk sensitivity in locomotion, enhancing safety in hazardous environments.

Contribution

It proposes Distributional Proximal Policy Optimization (DPPO), a new reinforcement learning approach that incorporates risk metrics for safer robot locomotion without extra reward tuning.

Findings

01

Emergent risk-sensitive behaviors in simulation

02

Successful real-world implementation on ANYmal robot

03

Adjustable risk preference parameter for behavior control

Abstract

Deployment in hazardous environments requires robots to understand the risks associated with their actions and movements to prevent accidents. Despite its importance, these risks are not explicitly modeled by currently deployed locomotion controllers for legged robots. In this work, we propose a risk sensitive locomotion training method employing distributional reinforcement learning to consider safety explicitly. Instead of relying on a value expectation, we estimate the complete value distribution to account for uncertainty in the robot's interaction with the environment. The value distribution is consumed by a risk metric to extract risk sensitive value estimates. These are integrated into Proximal Policy Optimization (PPO) to derive our method, Distributional Proximal Policy Optimization (DPPO). The risk preference, ranging from risk-averse to risk-seeking, can be controlled by a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.