Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?

Yujin Choi; Youngjoo Park; Junyoung Byun; Jaewook Lee; and Jinseong Park

arXiv:2505.22061·cs.CL·November 25, 2025

Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?

Yujin Choi, Youngjoo Park, Junyoung Byun, Jaewook Lee, and Jinseong Park

PDF

Open Access 1 Video

TL;DR

This paper presents a similarity-based detection framework for membership inference attacks in retrieval-augmented generation systems, effectively protecting private data while maintaining utility and system-agnosticism.

Contribution

It introduces a novel similarity-based detection method and a detect-and-hide strategy that obfuscates attackers without compromising system utility.

Findings

01

Successfully detects state-of-the-art MIAs

02

Effectively obfuscates attackers in RAG systems

03

Maintains data utility and system-agnosticism

Abstract

Retrieval-augmented generation (RAG) mitigates the hallucination problem in large language models (LLMs) and has proven effective for personalized usages. However, delivering private retrieved documents directly to LLMs introduces vulnerability to membership inference attacks (MIAs), which try to determine whether the target data point exists in the private external database or not. Based on the insight that MIA queries typically exhibit high similarity to only one target document, we introduce a novel similarity-based MIA detection framework designed for the RAG system. With the proposed method, we show that a simple detect-and-hide strategy can successfully obfuscate attackers, maintain data utility, and remain system-agnostic against MIA. We experimentally prove its detection and defense against various state-of-the-art MIA methods and its adaptability to existing RAG systems.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?· underline

Taxonomy

TopicsPrivacy-Preserving Technologies in Data · Cryptography and Data Security · Access Control and Trust