Improving Adversarial Robustness via Joint Classification and Multiple   Explicit Detection Classes

Sina Baharlouei; Fatemeh Sheikholeslami; Meisam Razaviyayn; Zico; Kolter

arXiv:2210.14410·cs.CV·May 12, 2023

Improving Adversarial Robustness via Joint Classification and Multiple Explicit Detection Classes

Sina Baharlouei, Fatemeh Sheikholeslami, Meisam Razaviyayn, Zico, Kolter

PDF

Open Access 1 Repo

TL;DR

This paper introduces a method to improve adversarial robustness in deep networks by extending joint classification and detection to multiple abstain classes, with regularization to prevent model degeneracy, resulting in better accuracy tradeoffs.

Contribution

It proposes a novel approach using multiple abstain classes with regularization to enhance certified robustness against adversarial attacks.

Findings

01

Outperforms state-of-the-art algorithms in robustness benchmarks.

02

Effectively balances standard and robust accuracy.

03

Regularization promotes full utilization of multiple abstain classes.

Abstract

This work concerns the development of deep networks that are certifiably robust to adversarial attacks. Joint robust classification-detection was recently introduced as a certified defense mechanism, where adversarial examples are either correctly classified or assigned to the "abstain" class. In this work, we show that such a provable framework can benefit by extension to networks with multiple explicit abstain classes, where the adversarial examples are adaptively assigned to those. We show that naively adding multiple abstain classes can lead to "model degeneracy", then we propose a regularization approach and a training method to counter this degeneracy by promoting full use of the multiple abstain classes. Our experiments demonstrate that the proposed approach consistently achieves favorable standard vs. robust verified accuracy tradeoffs, outperforming state-of-the-art algorithms…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

sinabaharlouei/multipleabstaindetection
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Anomaly Detection Techniques and Applications