Diffusion Imitation from Observation

Bo-Ruei Huang; Chun-Kai Yang; Chun-Mao Lai; Dai-Jie Wu; Shao-Hua Sun

arXiv:2410.05429·cs.LG·October 10, 2024

Diffusion Imitation from Observation

Bo-Ruei Huang, Chun-Kai Yang, Chun-Mao Lai, Dai-Jie Wu, Shao-Hua Sun

PDF

Open Access 1 Video

TL;DR

This paper introduces DIFO, a novel imitation learning framework that integrates diffusion models to generate state transitions from observations, improving imitation quality in continuous control tasks without action labels.

Contribution

The paper proposes a new diffusion-based approach for imitation from observation, replacing adversarial discriminators with diffusion models to enhance training stability and performance.

Findings

01

DIFO outperforms existing methods in multiple continuous control benchmarks.

02

The diffusion model effectively captures expert transition distributions.

03

The approach demonstrates robustness to hyperparameter variations.

Abstract

Learning from observation (LfO) aims to imitate experts by learning from state-only demonstrations without requiring action labels. Existing adversarial imitation learning approaches learn a generator agent policy to produce state transitions that are indistinguishable to a discriminator that learns to classify agent and expert state transitions. Despite its simplicity in formulation, these methods are often sensitive to hyperparameters and brittle to train. Motivated by the recent success of diffusion models in generative modeling, we propose to integrate a diffusion model into the adversarial imitation learning from observation framework. Specifically, we employ a diffusion model to capture expert and agent transitions by generating the next state, given the current state. Then, we reformulate the learning objective to train the diffusion model as a binary classifier and use it to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Diffusion Imitation from Observation· slideslive

Taxonomy

TopicsDigital Humanities and Scholarship · Natural Language Processing Techniques

MethodsDiffusion