A Collocation-based Method for Addressing Challenges in Word-level   Metric Differential Privacy

Stephen Meisenbacher; Maulik Chevli; and Florian Matthes

arXiv:2407.00638·cs.CL·July 2, 2024·1 cites

A Collocation-based Method for Addressing Challenges in Word-level Metric Differential Privacy

Stephen Meisenbacher, Maulik Chevli, and Florian Matthes

PDF

Open Access 1 Repo

TL;DR

This paper introduces a collocation-based approach to improve semantic coherence and flexibility in word-level metric differential privacy for NLP, addressing limitations of existing methods by perturbing n-grams instead of individual words.

Contribution

The authors propose a novel collocation-based method that operates between word and sentence levels, enhancing semantic coherence and output variability in differentially private NLP applications.

Findings

01

Improved semantic coherence in privatized text outputs.

02

Enhanced flexibility with variable length outputs.

03

Effective privacy preservation demonstrated through utility and privacy tests.

Abstract

Applications of Differential Privacy (DP) in NLP must distinguish between the syntactic level on which a proposed mechanism operates, often taking the form of $word-level$ or $document-level$ privatization. Recently, several word-level $Metric$ Differential Privacy approaches have been proposed, which rely on this generalized DP notion for operating in word embedding spaces. These approaches, however, often fail to produce semantically coherent textual outputs, and their application at the sentence- or document-level is only possible by a basic composition of word perturbations. In this work, we strive to address these challenges by operating $between$ the word and sentence levels, namely with $collocations$ . By perturbing n-grams rather than single words, we devise a method where composed privatized outputs have higher semantic coherence and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

sjmeis/CLMLDP
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAccess Control and Trust · Privacy-Preserving Technologies in Data