Regularizing Black-box Models for Improved Interpretability

Gregory Plumb; Maruan Al-Shedivat; Angel Alexander Cabrera; Adam; Perer; Eric Xing; Ameet Talwalkar

arXiv:1902.06787·cs.LG·November 10, 2020·37 cites

Regularizing Black-box Models for Improved Interpretability

Gregory Plumb, Maruan Al-Shedivat, Angel Alexander Cabrera, Adam, Perer, Eric Xing, Ameet Talwalkar

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces ExpO, a hybrid approach that regularizes black-box models during training to enhance explanation quality, resulting in more faithful and stable explanations without domain-specific knowledge.

Contribution

ExpO is a novel, differentiable, model-agnostic regularization method that improves post-hoc explanation quality for black-box models during training.

Findings

01

ExpO-regularized models produce explanations with higher fidelity.

02

ExpO improves explanation stability compared to baseline models.

03

User study confirms more useful explanations with ExpO.

Abstract

Most of the work on interpretable machine learning has focused on designing either inherently interpretable models, which typically trade-off accuracy for interpretability, or post-hoc explanation systems, whose explanation quality can be unpredictable. Our method, ExpO, is a hybridization of these approaches that regularizes a model for explanation quality at training time. Importantly, these regularizers are differentiable, model agnostic, and require no domain knowledge to define. We demonstrate that post-hoc explanations for ExpO-regularized models have better explanation quality, as measured by the common fidelity and stability metrics. We verify that improving these metrics leads to significantly more useful explanations with a user study on a realistic task.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

GDPlumb/ExpO
pytorchOfficial

Videos

Regularizing Black-box Models for Improved Interpretability· slideslive

Taxonomy

TopicsExplainable Artificial Intelligence (XAI) · Adversarial Robustness in Machine Learning · Machine Learning in Healthcare