Skill-aware Mutual Information Optimisation for Generalisation in   Reinforcement Learning

Xuehui Yu; Mhairi Dunion; Xin Li; Stefano V. Albrecht

arXiv:2406.04815·cs.LG·November 7, 2024

Skill-aware Mutual Information Optimisation for Generalisation in Reinforcement Learning

Xuehui Yu, Mhairi Dunion, Xin Li, Stefano V. Albrecht

PDF

Open Access 1 Repo

TL;DR

This paper introduces Skill-aware Mutual Information (SaMI) and SaNCE to improve the generalisation of Meta-Reinforcement Learning agents across tasks, especially with limited samples, by distinguishing skills in context embeddings.

Contribution

The paper proposes SaMI and SaNCE as novel methods to enhance task generalisation and robustness in Meta-RL, addressing the sample efficiency challenge.

Findings

01

SaMI improves zero-shot generalisation to unseen tasks.

02

SaNCE enhances robustness to fewer samples, mitigating the $ extlog$-$K$ curse.

03

Experimental results on MuJoCo and Panda-gym benchmarks validate effectiveness.

Abstract

Meta-Reinforcement Learning (Meta-RL) agents can struggle to operate across tasks with varying environmental features that require different optimal skills (i.e., different modes of behaviour). Using context encoders based on contrastive learning to enhance the generalisability of Meta-RL agents is now widely studied but faces challenges such as the requirement for a large sample size, also referred to as the $lo g$ - $K$ curse. To improve RL generalisation to different tasks, we first introduce Skill-aware Mutual Information (SaMI), an optimisation objective that aids in distinguishing context embeddings according to skills, thereby equipping RL agents with the ability to identify and execute different skills across tasks. We then propose Skill-aware Noise Contrastive Estimation (SaNCE), a $K$ -sample estimator used to optimise the SaMI objective. We provide a framework for equipping an…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

uoe-agents/sami
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics

MethodsContrastive Learning