Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents

Kevin Baum; Lisa Dargasz; Felix Jahn; Timo P. Gros; Verena Wolf

arXiv:2409.15014·cs.AI·June 17, 2025

Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents

Kevin Baum, Lisa Dargasz, Felix Jahn, Timo P. Gros, Verena Wolf

PDF

Open Access

TL;DR

This paper introduces a reinforcement learning framework that incorporates normative reasons to enable moral decision-making in artificial agents, using a reason-based shield generator and iterative improvement via moral feedback.

Contribution

It presents a novel architecture integrating normative reasons into reinforcement learning and an algorithm for refining moral decision-making through case-based feedback.

Findings

01

The reason-based shield effectively constrains agents to morally justified actions.

02

The iterative algorithm improves the moral reasoning of agents over time.

03

The approach aligns AI behavior with recognized normative principles.

Abstract

We propose an extension of the reinforcement learning architecture that enables moral decision-making of reinforcement learning agents based on normative reasons. Central to this approach is a reason-based shield generator yielding a moral shield that binds the agent to actions that conform with recognized normative reasons so that our overall architecture restricts the agent to actions that are (internally) morally justified. In addition, we describe an algorithm that allows to iteratively improve the reason-based shield generator through case-based feedback from a moral judge.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsPsychology of Moral and Emotional Judgment