Value Memory Graph: A Graph-Structured World Model for Offline   Reinforcement Learning

Deyao Zhu; Li Erran Li; Mohamed Elhoseiny

arXiv:2206.04384·cs.LG·May 3, 2023

Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning

Deyao Zhu, Li Erran Li, Mohamed Elhoseiny

PDF

Open Access 1 Repo 1 Video

TL;DR

The paper introduces Value Memory Graph (VMG), a discrete graph-based world model for offline RL that simplifies policy learning by abstracting environments and applying value iteration on a finite graph structure.

Contribution

The paper proposes VMG, a novel graph-structured world model for offline RL that enables efficient policy learning through value iteration on a simplified environment abstraction.

Findings

01

VMG outperforms state-of-the-art offline RL methods on D4RL benchmarks.

02

VMG is especially effective in environments with sparse rewards and long horizons.

03

The approach demonstrates the benefit of discrete graph models in complex RL tasks.

Abstract

Reinforcement Learning (RL) methods are typically applied directly in environments to learn policies. In some complex environments with continuous state-action spaces, sparse rewards, and/or long temporal horizons, learning a good policy in the original environments can be difficult. Focusing on the offline RL setting, we aim to build a simple and discrete world model that abstracts the original environment. RL methods are applied to our world model instead of the environment data for simplified policy learning. Our world model, dubbed Value Memory Graph (VMG), is designed as a directed-graph-based Markov decision process (MDP) of which vertices and directed edges represent graph states and graph actions, separately. As state-action spaces of VMG are finite and relatively small compared to the original environment, we can directly apply the value iteration algorithm on VMG to estimate…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

tsutikgiau/valuememorygraph
pytorchOfficial

Videos

Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement Learning· slideslive

Taxonomy

TopicsReinforcement Learning in Robotics