A Source-Criticism Debiasing Method for GloVe Embeddings

Hope McGovern

arXiv:2106.13382·cs.CL·June 28, 2021·1 cites

A Source-Criticism Debiasing Method for GloVe Embeddings

Hope McGovern

PDF

Open Access 1 Repo

TL;DR

This paper introduces SC-GloVe, a debiasing method for word embeddings that incorporates explicit bias information, reducing social biases without losing training data or performance.

Contribution

The paper proposes a novel source-criticism based debiasing technique for GloVe embeddings that preserves data and accuracy while reducing biases.

Findings

01

Reduces bias effect size on WEAT tests

02

Maintains training set size and TOP-1 performance

03

Runs efficiently with a bias gradient approximation

Abstract

It is well-documented that word embeddings trained on large public corpora consistently exhibit known human social biases. Although many methods for debiasing exist, almost all fixate on completely eliminating biased information from the embeddings and often diminish training set size in the process. In this paper, we present a simple yet effective method for debiasing GloVe word embeddings (Pennington et al., 2014) which works by incorporating explicit information about training set bias rather than removing biased data outright. Our method runs quickly and efficiently with the help of a fast bias gradient approximation method from Brunet et al. (2019). As our approach is akin to the notion of 'source criticism' in the humanities, we term our method Source-Critical GloVe (SC-GloVe). We show that SC-GloVe reduces the effect size on Word Embedding Association Test (WEAT) sets without…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

mebrunet/understanding-bias
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Computational and Text Analysis Methods

MethodsGloVe Embeddings