Unintended Impacts of LLM Alignment on Global Representation

Michael J. Ryan; William Held; Diyi Yang

arXiv:2402.15018·cs.CL·June 10, 2024·3 cites

Unintended Impacts of LLM Alignment on Global Representation

Michael J. Ryan, William Held, Diyi Yang

PDF

Open Access 1 Repo 1 Datasets

TL;DR

This paper investigates how aligning Large Language Models to human preferences can unintentionally cause disparities in global representation, affecting dialects, multilingualism, and international opinions, and offers recommendations for more equitable tuning.

Contribution

It reveals unintended disparities caused by current alignment procedures and discusses design choices for more equitable preference tuning in LLMs.

Findings

01

Alignment creates disparities between English dialects and global opinions.

02

Alignment improves capabilities in several languages.

03

Unintended impacts highlight need for more equitable tuning strategies.

Abstract

Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

salt-nlp/unintended-impacts-of-alignment
pytorchOfficial

Datasets

SALT-NLP/AskRedditCountries
dataset· 16 dl
16 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsComparative and International Law Studies · Border Security and International Relations · Dispute Resolution and Class Actions

MethodsFocus · ALIGN