Attacks Which Do Not Kill Training Make Adversarial Learning Stronger

Jingfeng Zhang; Xilie Xu; Bo Han; Gang Niu; Lizhen Cui; Masashi; Sugiyama; Mohan Kankanhalli

arXiv:2002.11242·cs.LG·September 8, 2020·84 cites

Attacks Which Do Not Kill Training Make Adversarial Learning Stronger

Jingfeng Zhang, Xilie Xu, Bo Han, Gang Niu, Lizhen Cui, Masashi, Sugiyama, Mohan Kankanhalli

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces friendly adversarial training (FAT), a novel method that improves adversarial robustness without sacrificing natural generalization by using less adversarial data through early stopping of PGD.

Contribution

The paper proposes FAT, a new adversarial training approach that employs less adversarial data, justified theoretically and validated empirically, challenging the trade-off between robustness and generalization.

Findings

01

FAT achieves robustness without harming natural accuracy.

02

Early-stopped PGD simplifies adversarial data generation.

03

Theoretical bounds support FAT's effectiveness.

Abstract

Adversarial training based on the minimax formulation is necessary for obtaining adversarial robustness of trained models. However, it is conservative or even pessimistic so that it sometimes hurts the natural generalization. In this paper, we raise a fundamental question---do we have to trade off natural generalization for adversarial robustness? We argue that adversarial training is to employ confident adversarial data for updating the current model. We propose a novel approach of friendly adversarial training (FAT): rather than employing most adversarial data maximizing the loss, we search for least adversarial (i.e., friendly adversarial) data minimizing the loss, among the adversarial data that are confidently misclassified. Our novel formulation is easy to implement by just stopping the most adversarial data searching algorithms such as PGD (projected gradient descent) early,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

zjfheart/Friendly-Adversarial-Training
pytorch

Videos

Attacks Which Do Not Kill Training Make Adversarial Learning Stronger· slideslive

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Domain Adaptation and Few-Shot Learning · Anomaly Detection Techniques and Applications