Learning on a Budget via Teacher Imitation

Ercument Ilhan; Jeremy Gow; Diego Perez-Liebana

arXiv:2104.08440·cs.LG·July 1, 2021

Learning on a Budget via Teacher Imitation

Ercument Ilhan, Jeremy Gow, Diego Perez-Liebana

PDF

1 Repo

TL;DR

This paper introduces a unified, adaptive approach to learning from teacher advice in reinforcement learning, optimizing advice collection and utilization under budget constraints, and demonstrating strong performance in Atari games.

Contribution

It extends teacher imitation to unify advice collection and use, with automatic hyperparameter tuning for broad applicability and simplicity.

Findings

01

Outperforms or matches top competitors in Atari games

02

Components provide significant individual advantages

03

Automatically adapts to different tasks with minimal human intervention

Abstract

Deep Reinforcement Learning (RL) techniques can benefit greatly from leveraging prior experience, which can be either self-generated or acquired from other entities. Action advising is a framework that provides a flexible way to transfer such knowledge in the form of actions between teacher-student peers. However, due to the realistic concerns, the number of these interactions is limited with a budget; therefore, it is crucial to perform these in the most appropriate moments. There have been several promising studies recently that address this problem setting especially from the student's perspective. Despite their success, they have some shortcomings when it comes to the practical applicability and integrity as an overall solution to the learning from advice challenge. In this paper, we extend the idea of advice reusing via teacher imitation to construct a unified approach that…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ercumentilhan/advice-imitation-reuse
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.