The Effects of Approximate Multiplication on Convolutional Neural   Networks

Min Soo Kim; Alberto A. Del Barrio; HyunJin Kim; Nader Bagherzadeh

arXiv:2007.10500·cs.LG·January 12, 2021

The Effects of Approximate Multiplication on Convolutional Neural Networks

Min Soo Kim, Alberto A. Del Barrio, HyunJin Kim, Nader Bagherzadeh

PDF

1 Repo

TL;DR

This paper investigates how approximate multiplication affects CNN inference accuracy and efficiency, showing that certain approximations can significantly reduce energy consumption with minimal accuracy loss.

Contribution

It provides an analytical explanation for why approximate multipliers like Mitch-$w$6 maintain accuracy and demonstrates their effectiveness in CNN inference without retraining.

Findings

01

Approximate multipliers can achieve near-FP32 accuracy in CNNs.

02

Mitch-$w$6 reduces energy consumption by up to 80% compared to bfloat16.

03

Analytical justification that multiplications can be approximated while additions remain exact.

Abstract

This paper analyzes the effects of approximate multiplication when performing inferences on deep convolutional neural networks (CNNs). The approximate multiplication can reduce the cost of the underlying circuits so that CNN inferences can be performed more efficiently in hardware accelerators. The study identifies the critical factors in the convolution, fully-connected, and batch normalization layers that allow more accurate CNN predictions despite the errors from approximate multiplication. The same factors also provide an arithmetic explanation of why bfloat16 multiplication performs well on CNNs. The experiments are performed with recognized network architectures to show that the approximate multipliers can produce predictions that are nearly as accurate as the FP32 references, without additional training. For example, the ResNet and Inception-v4 models with Mitch- $w$ 6…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

albertodbg/log-arithmetic
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

Methods*Communicated@Fast*How Do I Communicate to Expedia? · Kaiming Initialization · Residual Block · Residual Connection · Inception-B · Softmax · Bottleneck Residual Block · Max Pooling · 1x1 Convolution · Global Average Pooling