Interpreting Adversarial Examples by Activation Promotion and   Suppression

Kaidi Xu; Sijia Liu; Gaoyuan Zhang; Mengshu Sun; Pu Zhao; Quanfu Fan,; Chuang Gan; Xue Lin

arXiv:1904.02057·cs.CV·October 2, 2019·39 cites

Interpreting Adversarial Examples by Activation Promotion and Suppression

Kaidi Xu, Sijia Liu, Gaoyuan Zhang, Mengshu Sun, Pu Zhao, Quanfu Fan,, Chuang Gan, Xue Lin

PDF

Open Access

TL;DR

This paper investigates how adversarial perturbations affect CNN activations, categorizing them into promotion, suppression, or balanced types, and links these effects to interpretability and robustness improvements.

Contribution

It introduces a promotion-suppression effect framework for understanding adversarial perturbations and connects pixel-level effects to class-specific regions and concept-level interpretability.

Findings

01

Adversarial perturbations can be categorized into suppression, promotion, or balanced types.

02

Pixel-level perturbations relate to class-specific discriminative regions via CAM.

03

Sensitivity of units to attacks correlates with their semantic interpretability.

Abstract

It is widely known that convolutional neural networks (CNNs) are vulnerable to adversarial examples: images with imperceptible perturbations crafted to fool classifiers. However, interpretability of these perturbations is less explored in the literature. This work aims to better understand the roles of adversarial perturbations and provide visual explanations from pixel, image and network perspectives. We show that adversaries have a promotion-suppression effect (PSE) on neurons' activations and can be primarily categorized into three types: i) suppression-dominated perturbations that mainly reduce the classification score of the true label, ii) promotion-dominated perturbations that focus on boosting the confidence of the target label, and iii) balanced perturbations that play a dual role in suppression and promotion. We also provide image-level interpretability of adversarial…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Explainable Artificial Intelligence (XAI) · Integrated Circuits and Semiconductor Failure Analysis

MethodsNetwork Dissection · Interpretability