MINIMAL: Mining Models for Data Free Universal Adversarial Triggers

Swapnil Parekh; Yaman Singla Kumar; Somesh Singh; Changyou Chen,; Balaji Krishnamurthy; and Rajiv Ratn Shah

arXiv:2109.12406·cs.CL·September 28, 2021

MINIMAL: Mining Models for Data Free Universal Adversarial Triggers

Swapnil Parekh, Yaman Singla Kumar, Somesh Singh, Changyou Chen,, Balaji Krishnamurthy, and Rajiv Ratn Shah

PDF

Open Access 1 Repo

TL;DR

MINIMAL introduces a novel data-free method to mine universal adversarial triggers for NLP models, achieving significant accuracy drops comparable to data-dependent approaches, thus exposing vulnerabilities without requiring large datasets.

Contribution

The paper presents MINIMAL, a data-free algorithm for mining universal adversarial triggers, reducing reliance on large datasets for generating effective adversarial inputs.

Findings

01

Achieves 93.6% to 9.6% accuracy drop on SST positive class.

02

Reduces SNLI entailment accuracy from 90.95% to below 0.6%.

03

Matches the effectiveness of data-dependent methods without using data.

Abstract

It is well known that natural language models are vulnerable to adversarial attacks, which are mostly input-specific in nature. Recently, it has been shown that there also exist input-agnostic attacks in NLP models, called universal adversarial triggers. However, existing methods to craft universal triggers are data intensive. They require large amounts of data samples to generate adversarial triggers, which are typically inaccessible by attackers. For instance, previous works take 3000 data samples per class for the SNLI dataset to generate adversarial triggers. In this paper, we present a novel data-free approach, MINIMAL, to mine input-agnostic adversarial triggers from models. Using the triggers produced with our data-free algorithm, we reduce the accuracy of Stanford Sentiment Treebank's positive class from 93.6% to 9.6%. Similarly, for the Stanford Natural Language Inference…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

midas-research/data-free-uats
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Advanced Malware Detection Techniques · Anomaly Detection Techniques and Applications