Towards the Anonymization of the Language Modeling

Antoine Boutet; Lucas Magnana; Juliette S\'en\'echal

arXiv:2501.02407·cs.CL·May 21, 2026

Towards the Anonymization of the Language Modeling

Antoine Boutet, Lucas Magnana, Juliette S\'en\'echal

PDF

TL;DR

This paper introduces privacy-preserving methods for language model anonymization, using masking and causal modeling techniques to prevent memorization of sensitive data while maintaining utility.

Contribution

It proposes novel MLM and CLM approaches for anonymizing language models, specifically tailored to protect sensitive information in medical datasets.

Findings

01

Models effectively prevent memorization of personal identifiers.

02

Proposed methods maintain high utility while enhancing privacy.

03

Evaluation on medical data shows promising privacy-utility tradeoffs.

Abstract

Rapid advances in Natural Language Processing (NLP) have revolutionized many fields, including healthcare. However, these advances raise significant privacy concerns, especially when pre-trained models fine-tuned and specialized on sensitive data can memorize and then expose and regurgitate personal information. This paper presents a privacy-preserving language modeling approach to address the problem of language models anonymization, and thus promote their sharing. Specifically, we propose both a Masking Language Modeling (MLM) methodology to specialize a BERT-like language model, and a Causal Language Modeling (CLM) methodology to specialize a GPT-like model that avoids the model from memorizing direct and indirect identifying information present in the training data. We have comprehensively evaluated our approaches using a medical dataset and compared them against different…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Quality and Management