A Geometry-Inspired Attack for Generating Natural Language Adversarial   Examples

Zhao Meng; Roger Wattenhofer

arXiv:2010.01345·cs.CL·October 6, 2020

A Geometry-Inspired Attack for Generating Natural Language Adversarial Examples

Zhao Meng, Roger Wattenhofer

PDF

TL;DR

This paper introduces a geometry-inspired method for creating natural language adversarial examples that effectively fool models with minimal word changes, maintaining human indistinguishability and enhancing model robustness.

Contribution

The paper presents a novel geometry-inspired attack technique that approximates decision boundaries to generate effective natural language adversarial examples.

Findings

01

High success rate in fooling models with few word replacements

02

Adversarial examples are hard for humans to recognize

03

Adversarial training improves model robustness

Abstract

Generating adversarial examples for natural language is hard, as natural language consists of discrete symbols, and examples are often of variable lengths. In this paper, we propose a geometry-inspired attack for generating natural language adversarial examples. Our attack generates adversarial examples by iteratively approximating the decision boundary of Deep Neural Networks (DNNs). Experiments on two datasets with two different models show that our attack fools natural language models with high success rates, while only replacing a few words. Human evaluation shows that adversarial examples generated by our attack are hard for humans to recognize. Further experiments show that adversarial training can improve model robustness against our attack.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.