OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Yulin Hu; Zimo Long; Jiahe Guo; Xingyu Sui; Xing Fu; Weixiang Zhao; Yanyan Zhao; Bing Qin

arXiv:2601.13722·cs.CL·January 21, 2026

OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Yulin Hu, Zimo Long, Jiahe Guo, Xingyu Sui, Xing Fu, Weixiang Zhao, Yanyan Zhao, Bing Qin

PDF

Open Access

TL;DR

This paper introduces OP-Bench, a benchmark for evaluating over-personalization in memory-augmented conversational agents, highlighting its prevalence and proposing a filtering method to improve appropriateness.

Contribution

We formalize over-personalization into three types, create OP-Bench with 1,700 instances, and propose Self-ReCheck to mitigate over-personalization in dialogue systems.

Findings

01

Over-personalization is widespread in memory-augmented models.

02

Agents often retrieve unnecessary user memories.

03

Self-ReCheck reduces over-personalization while maintaining personalization.

Abstract

Memory-augmented conversational agents enable personalized interactions using long-term user memory and have gained substantial traction. However, existing benchmarks primarily focus on whether agents can recall and apply user information, while overlooking whether such personalization is used appropriately. In fact, agents may overuse personal information, producing responses that feel forced, intrusive, or socially inappropriate to users. We refer to this issue as \emph{over-personalization}. In this work, we formalize over-personalization into three types: Irrelevance, Repetition, and Sycophancy, and introduce \textbf{OP-Bench} a benchmark of 1,700 verified instances constructed from long-horizon dialogue histories. Using \textbf{OP-Bench}, we evaluate multiple large language models and memory-augmentation methods, and find that over-personalization is widespread when memory is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAI in Service Interactions · Topic Modeling · Social Robot Interaction and HRI