Adaptive Pairwise Weights for Temporal Credit Assignment

Zeyu Zheng; Risto Vuorio; Richard Lewis; Satinder Singh

arXiv:2102.04999·cs.LG·June 7, 2022

Adaptive Pairwise Weights for Temporal Credit Assignment

Zeyu Zheng, Risto Vuorio, Richard Lewis, Satinder Singh

PDF

Open Access 1 Video

TL;DR

This paper introduces a learned pairwise weighting scheme for temporal credit assignment in reinforcement learning, improving performance over traditional methods by dynamically adapting weights during training.

Contribution

It proposes a novel metagradient approach to learn pairwise weight functions for credit assignment, surpassing fixed heuristics in RL tasks.

Findings

01

Learned pairwise weights often outperform fixed heuristics.

02

Dynamic weighting improves RL policy performance.

03

Method adapts weights during training for better credit assignment.

Abstract

How much credit (or blame) should an action taken in a state get for a future reward? This is the fundamental temporal credit assignment problem in Reinforcement Learning (RL). One of the earliest and still most widely used heuristics is to assign this credit based on a scalar coefficient, $λ$ (treated as a hyperparameter), raised to the power of the time interval between the state-action and the reward. In this empirical paper, we explore heuristics based on more general pairwise weightings that are functions of the state in which the action was taken, the state at the time of the reward, as well as the time interval between the two. Of course it isn't clear what these pairwise weight functions should be, and because they are too complex to be treated as hyperparameters we develop a metagradient procedure for learning these weight functions during the usual RL training of a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Adaptive Pairwise Weights for Temporal Credit Assignment· underline

Taxonomy

TopicsReinforcement Learning in Robotics