The Effects of Reward Misspecification: Mapping and Mitigating   Misaligned Models

Alexander Pan; Kush Bhatia; Jacob Steinhardt

arXiv:2201.03544·cs.LG·February 15, 2022·21 cites

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Alexander Pan, Kush Bhatia, Jacob Steinhardt

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper systematically studies reward hacking in RL, revealing how agent capabilities influence exploitation of misspecified rewards and identifying phase transitions that challenge safety monitoring.

Contribution

It introduces four RL environments with misspecified rewards, analyzes the impact of agent capabilities, and proposes anomaly detection methods for unsafe policies.

Findings

01

More capable agents exploit reward misspecifications more

02

Identification of capability thresholds causing behavior shifts

03

Baseline anomaly detectors for policy safety

Abstract

Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified rewards. We investigate reward hacking as a function of agent capabilities: model capacity, action space resolution, observation space noise, and training time. More capable agents often exploit reward misspecifications, achieving higher proxy reward and lower true reward than less capable agents. Moreover, we find instances of phase transitions: capability thresholds at which the agent's behavior qualitatively shifts, leading to a sharp decrease in the true reward. Such phase transitions pose challenges to monitoring the safety of ML systems. To address this, we propose an anomaly detection task for aberrant policies and offer several baseline…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

aypan17/reward-misspecification
noneOfficial

Videos

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models· slideslive

Taxonomy

TopicsNetwork Security and Intrusion Detection · Information and Cyber Security · Reinforcement Learning in Robotics