On Network Science and Mutual Information for Explaining Deep Neural   Networks

Brian Davis; Umang Bhatt; Kartikeya Bhardwaj; Radu Marculescu; Jos\'e; M. F. Moura

arXiv:1901.08557·cs.LG·May 5, 2020

On Network Science and Mutual Information for Explaining Deep Neural Networks

Brian Davis, Umang Bhatt, Kartikeya Bhardwaj, Radu Marculescu, Jos\'e, M. F. Moura

PDF

TL;DR

This paper introduces NIF, a novel method combining mutual information and network science to interpret deep neural networks by quantifying information flow between neurons.

Contribution

It presents a new approach, NIF, that approximates mutual information to analyze and explain the internal information dynamics of deep learning models.

Findings

01

NIF effectively quantifies information flow between neurons.

02

The method exposes internal model internals and aids feature attribution.

03

Provides a new perspective on deep learning interpretability.

Abstract

In this paper, we present a new approach to interpret deep learning models. By coupling mutual information with network science, we explore how information flows through feedforward networks. We show that efficiently approximating mutual information allows us to create an information measure that quantifies how much information flows between any two neurons of a deep learning model. To that end, we propose NIF, Neural Information Flow, a technique for codifying information flow that exposes deep learning model internals and provides feature attributions.

Figures11

Click any figure to enlarge with its caption.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.