Investigating Human-Identifiable Features Hidden in Adversarial   Perturbations

Dennis Y. Menn; Tzu-hsun Feng; Sriram Vishwanath; Hung-yi Lee

arXiv:2309.16878·cs.LG·October 2, 2023

Investigating Human-Identifiable Features Hidden in Adversarial Perturbations

Dennis Y. Menn, Tzu-hsun Feng, Sriram Vishwanath, Hung-yi Lee

PDF

Open Access

TL;DR

This paper investigates human-identifiable features in adversarial perturbations across multiple attack algorithms and datasets, revealing effects like masking and generation, and providing insights into transferability and model interpretability.

Contribution

It uncovers human-identifiable features in adversarial perturbations and distinguishes effects in targeted versus untargeted attacks, advancing understanding of attack mechanisms.

Findings

01

Identification of human-identifiable features in adversarial perturbations

02

Distinct masking and generation effects in untargeted and targeted attacks

03

Perturbations show similarity across different attack algorithms and models

Abstract

Neural networks perform exceedingly well across various machine learning tasks but are not immune to adversarial perturbations. This vulnerability has implications for real-world applications. While much research has been conducted, the underlying reasons why neural networks fall prey to adversarial attacks are not yet fully understood. Central to our study, which explores up to five attack algorithms across three datasets, is the identification of human-identifiable features in adversarial perturbations. Additionally, we uncover two distinct effects manifesting within human-identifiable features. Specifically, the masking effect is prominent in untargeted attacks, while the generation effect is more common in targeted attacks. Using pixel-level annotations, we extract such features and demonstrate their ability to compromise target models. In addition, our findings indicate a notable…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Explainable Artificial Intelligence (XAI) · Machine Learning in Materials Science