Hierarchical Variational Imitation Learning of Control Programs

Roy Fox; Richard Shin; William Paul; Yitian Zou; Dawn Song; Ken; Goldberg; Pieter Abbeel; Ion Stoica

arXiv:1912.12612·cs.LG·January 1, 2020·1 cites

Hierarchical Variational Imitation Learning of Control Programs

Roy Fox, Richard Shin, William Paul, Yitian Zou, Dawn Song, Ken, Goldberg, Pieter Abbeel, Ion Stoica

PDF

Open Access 1 Repo

TL;DR

This paper introduces a variational inference approach for hierarchical imitation learning, enabling autonomous agents to learn structured control programs more efficiently and accurately from demonstrations.

Contribution

It proposes a novel variational inference method for learning hierarchical control policies represented as program-like structures, improving data efficiency and generalization.

Findings

01

Outperforms LSTM baselines in data efficiency and accuracy

02

Achieves 24% error in bubble sort task with less data

03

Executes Karel programs flawlessly

Abstract

Autonomous agents can learn by imitating teacher demonstrations of the intended behavior. Hierarchical control policies are ubiquitously useful for such learning, having the potential to break down structured tasks into simpler sub-tasks, thereby improving data efficiency and generalization. In this paper, we propose a variational inference method for imitation learning of a control policy represented by parametrized hierarchical procedures (PHP), a program-like structure in which procedures can invoke sub-procedures to perform sub-tasks. Our method discovers the hierarchical structure in a dataset of observation-action traces of teacher demonstrations, by learning an approximate posterior distribution over the latent sequence of procedure calls and terminations. Samples from this learned distribution then guide the training of the hierarchical control policy. We identify and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

royf/hvil
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Robot Manipulation and Learning · Machine Learning and Algorithms

MethodsSigmoid Activation · Tanh Activation · Long Short-Term Memory