Theory-Grounded Measurement of U.S. Social Stereotypes in English   Language Models

Yang Trista Cao; Anna Sotnikova; Hal Daum\'e III; Rachel Rudinger,; Linda Zou

arXiv:2206.11684·cs.CL·June 24, 2022

Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

Yang Trista Cao, Anna Sotnikova, Hal Daum\'e III, Rachel Rudinger,, Linda Zou

PDF

Open Access 1 Repo

TL;DR

This paper introduces a new framework and measurement tool based on social psychology to systematically identify and analyze stereotypes in English language models, including intersectional biases.

Contribution

It adapts the ABC stereotype model for NLP, develops the sensitivity test (SeT), and evaluates stereotypes against human judgments, extending to intersectional identities.

Findings

01

SeT effectively measures stereotypic associations in LMs.

02

Comparison with human judgments validates the framework.

03

Framework captures intersectional stereotypes in language models.

Abstract

NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of Koch et al. (2016) from social psychology as a framework for the systematic study and discovery of stereotypic group-trait associations in language models (LMs). We introduce the sensitivity test (SeT) for measuring stereotypical associations from language models. To evaluate SeT and other measures using the ABC model, we collect group-trait judgments from U.S.-based subjects to compare with English LM stereotypes. Finally, we extend this framework to measure LM stereotyping of intersectional identities.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

tristacao/u.s_stereotypes
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHate Speech and Cyberbullying Detection · Migration, Health and Trauma · Computational and Text Analysis Methods

MethodsTest · Approximate Bayesian Computation