Output Randomization: A Novel Defense for both White-box and Black-box   Adversarial Models

Daniel Park; Haidar Khan; Azer Khan; Alex Gittens; B\"ulent Yener

arXiv:2107.03806·cs.LG·July 9, 2021·1 cites

Output Randomization: A Novel Defense for both White-box and Black-box Adversarial Models

Daniel Park, Haidar Khan, Azer Khan, Alex Gittens, B\"ulent Yener

PDF

Open Access

TL;DR

This paper introduces output randomization as a novel, effective defense mechanism against both white-box and black-box adversarial attacks on neural networks, reducing attack success rates significantly.

Contribution

It proposes two output randomization-based defenses: one at test time for black-box attacks and another during training for white-box attacks, both overcoming limitations of prior methods.

Findings

01

Black box attack success rate reduced to 0% with output randomization.

02

White box attack success rate reduced to 12% with output randomization training.

03

Defense is low overhead and compatible with various architectures.

Abstract

Adversarial examples pose a threat to deep neural network models in a variety of scenarios, from settings where the adversary has complete knowledge of the model in a "white box" setting and to the opposite in a "black box" setting. In this paper, we explore the use of output randomization as a defense against attacks in both the black box and white box models and propose two defenses. In the first defense, we propose output randomization at test time to thwart finite difference attacks in black box settings. Since this type of attack relies on repeated queries to the model to estimate gradients, we investigate the use of randomization to thwart such adversaries from successfully creating adversarial examples. We empirically show that this defense can limit the success rate of a black box adversary using the Zeroth Order Optimization attack to 0%. Secondly, we propose output…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Physical Unclonable Functions (PUFs) and Hardware Security · Integrated Circuits and Semiconductor Failure Analysis