Role-Playing Evaluation for Large Language Models

Yassine El Boudouri; Walter Nuninger; Julian Alvarez; Yvan Peter

arXiv:2505.13157·cs.CL·May 20, 2025

Role-Playing Evaluation for Large Language Models

Yassine El Boudouri, Walter Nuninger, Julian Alvarez, Yvan Peter

PDF

Open Access 1 Repo 2 Datasets

TL;DR

This paper introduces RPEval, a new benchmark for evaluating large language models' role-playing abilities across emotional, decision-making, moral, and consistency dimensions, addressing evaluation challenges.

Contribution

The paper presents RPEval, a comprehensive benchmark for assessing LLM role-playing skills, including its construction and baseline evaluations.

Findings

01

RPEval effectively measures LLM role-playing capabilities.

02

Baseline results highlight strengths and weaknesses of current models.

03

The benchmark is publicly available for further research.

Abstract

Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However, evaluating this ability presents significant challenges, as human assessments are resource-intensive and automated evaluations can be biased. To address this, we introduce Role-Playing Eval (RPEval), a novel benchmark designed to assess LLM role-playing capabilities across four key dimensions: emotional understanding, decision-making, moral alignment, and in-character consistency. This article details the construction of RPEval and presents baseline evaluations. Our code and dataset are available at https://github.com/yelboudouri/RPEval

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yelboudouri/rpeval
noneOfficial

Datasets

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsArtificial Intelligence in Healthcare and Education · Multimodal Machine Learning Applications · Machine Learning in Healthcare