High-level Modeling of Manufacturing Faults in Deep Neural Network   Accelerators

Shamik Kundu; Ahmet Soyyi\u{g}it; Khaza Anuarul Hoque; Kanad Basu

arXiv:2006.03616·cs.LG·October 27, 2020

High-level Modeling of Manufacturing Faults in Deep Neural Network Accelerators

Shamik Kundu, Ahmet Soyyi\u{g}it, Khaza Anuarul Hoque, Kanad Basu

PDF

TL;DR

This paper presents a formal probabilistic model of manufacturing faults in TPU neural network accelerators, analyzing their impact on classification accuracy through model checking and experimental validation.

Contribution

It introduces a formal DTMC-based model for permanent faults in TPU hardware and analyzes their effect on DNN inference accuracy.

Findings

01

Fault type and location significantly affect accuracy

02

Model checking quantifies fault impact probabilities

03

Experimental validation confirms theoretical predictions

Abstract

The advent of data-driven real-time applications requires the implementation of Deep Neural Networks (DNNs) on Machine Learning accelerators. Google's Tensor Processing Unit (TPU) is one such neural network accelerator that uses systolic array-based matrix multiplication hardware for computation in its crux. Manufacturing faults at any state element of the matrix multiplication unit can cause unexpected errors in these inference networks. In this paper, we propose a formal model of permanent faults and their propagation in a TPU using the Discrete-Time Markov Chain (DTMC) formalism. The proposed model is analyzed using the probabilistic model checking technique to reason about the likelihood of faulty outputs. The obtained quantitative results show that the classification accuracy is sensitive to the type of permanent faults as well as their location, bit position and the number of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.