Hierarchical Deep Q-Network from Imperfect Demonstrations in Minecraft

Alexey Skrynnik; Aleksey Staroverov; Ermek Aitygulov; Kirill Aksenov,; Vasilii Davydov; Aleksandr I. Panov

arXiv:1912.08664·cs.AI·July 14, 2020

Hierarchical Deep Q-Network from Imperfect Demonstrations in Minecraft

Alexey Skrynnik, Aleksey Staroverov, Ermek Aitygulov, Kirill Aksenov,, Vasilii Davydov, Aleksandr I. Panov

PDF

1 Repo

TL;DR

This paper introduces a Hierarchical Deep Q-Network that effectively learns from imperfect demonstrations in Minecraft, utilizing hierarchical structures, meta-actions, and adaptive replay buffers to improve performance.

Contribution

The paper proposes a novel HDQfD algorithm that handles imperfect demonstrations and extracts hierarchical structures from expert trajectories in Minecraft.

Findings

01

HDQfD achieved first place in the MineRL competition.

02

The method effectively filters poor-quality demonstration data.

03

Hierarchical structure improves learning efficiency.

Abstract

We present Hierarchical Deep Q-Network (HDQfD) that took first place in the MineRL competition. HDQfD works on imperfect demonstrations and utilizes the hierarchical structure of expert trajectories. We introduce the procedure of extracting an effective sequence of meta-actions and subgoals from demonstration data. We present a structured task-dependent replay buffer and adaptive prioritizing technique that allow the HDQfD agent to gradually erase poor-quality expert data from the buffer. In this paper, we present the details of the HDQfD algorithm and give the experimental results in the Minecraft domain.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

cog-isa/forger
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.