CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement

Gauri Kholkar; Ratinder Ahuja

arXiv:2505.12368·cs.CL·June 18, 2025

CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement

Gauri Kholkar, Ratinder Ahuja

PDF

Open Access 1 Video

TL;DR

CAPTURE introduces a new context-aware benchmark for prompt injection, revealing current guardrails' limitations and demonstrating a trained model that significantly improves detection accuracy and robustness.

Contribution

The paper presents CAPTURE, a novel benchmark for context-aware prompt injection testing, and introduces CaptureGuard, a model that enhances detection and reduces false positives and negatives.

Findings

01

Current guardrails have high false negatives in adversarial cases.

02

Existing models suffer from over-defense, causing false positives.

03

CaptureGuard outperforms existing defenses on context-aware datasets.

Abstract

Prompt injection remains a major security risk for large language models. However, the efficacy of existing guardrail models in context-aware settings remains underexplored, as they often rely on static attack benchmarks. Additionally, they have over-defense tendencies. We introduce CAPTURE, a novel context-aware benchmark assessing both attack detection and over-defense tendencies with minimal in-domain examples. Our experiments reveal that current prompt injection guardrail models suffer from high false negatives in adversarial cases and excessive false positives in benign scenarios, highlighting critical limitations. To demonstrate our framework's utility, we train CaptureGuard on our generated data. This new model drastically reduces both false negative and false positive rates on our context-aware datasets while also generalizing effectively to external benchmarks, establishing a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement· underline

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Topic Modeling · Security and Verification in Computing