Latent Space Explanation by Intervention

Itai Gat; Guy Lorberbom; Idan Schwartz; Tamir Hazan

arXiv:2112.04895·cs.LG·December 10, 2021

Latent Space Explanation by Intervention

Itai Gat, Guy Lorberbom, Idan Schwartz, Tamir Hazan

PDF

Open Access 1 Video

TL;DR

This paper introduces an intervention-based method using discrete variational autoencoders to interpret and visualize hidden concepts in neural networks, revealing biases and mechanisms behind predictions.

Contribution

It proposes a novel approach for explaining neural network decisions by intervening in the latent space and visualizing the effects, enhancing interpretability.

Findings

01

Effectively visualized biases in CelebA data

02

Identified concepts that influence class predictions

03

Demonstrated intervention can alter model biases

Abstract

The success of deep neural nets heavily relies on their ability to encode complex relations between their input and their output. While this property serves to fit the training data well, it also obscures the mechanism that drives prediction. This study aims to reveal hidden concepts by employing an intervention mechanism that shifts the predicted class based on discrete variational autoencoders. An explanatory model then visualizes the encoded information from any hidden layer and its corresponding intervened representation. By the assessment of differences between the original representation and the intervened representation, one can determine the concepts that can alter the class, hence providing interpretability. We demonstrate the effectiveness of our approach on CelebA, where we show various visualizations for bias in the data and suggest different interventions to reveal and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Latent Space Explanation by Intervention· underline

Taxonomy

TopicsExplainable Artificial Intelligence (XAI) · Anomaly Detection Techniques and Applications · Machine Learning and Data Classification