Do GPT Language Models Suffer From Split Personality Disorder? The   Advent Of Substrate-Free Psychometrics

Peter Romero; Stephen Fitz; Teruo Nakatsuma

arXiv:2408.07377·cs.CL·August 16, 2024

Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics

Peter Romero, Stephen Fitz, Teruo Nakatsuma

PDF

TL;DR

This paper investigates the stability of personality traits in large language models across languages, revealing inconsistencies that could impact AI safety, and proposes a new substrate-free psychometric framework.

Contribution

It introduces a novel substrate-free psychometric approach and demonstrates the instability of personality traits in language models across languages.

Findings

01

Language models show interlingual and intralingual personality instability.

02

Current models lack a consistent core personality, risking unsafe AI behavior.

03

Bayesian analysis reveals deeper-rooted issues in model personality traits.

Abstract

Previous research on emergence in large language models shows these display apparent human-like abilities and psychological latent traits. However, results are partly contradicting in expression and magnitude of these latent traits, yet agree on the worrisome tendencies to score high on the Dark Triad of narcissism, psychopathy, and Machiavellianism, which, together with a track record of derailments, demands more rigorous research on safety of these models. We provided a state of the art language model with the same personality questionnaire in nine languages, and performed Bayesian analysis of Gaussian Mixture Model, finding evidence for a deeper-rooted issue. Our results suggest both interlingual and intralingual instabilities, which indicate that current language models do not develop a consistent core personality. This can lead to unsafe behaviour of artificial intelligence systems…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.