WRPN & Apprentice: Methods for Training and Inference using   Low-Precision Numerics

Asit Mishra; Debbie Marr

arXiv:1803.00227·cs.CV·March 2, 2018

WRPN & Apprentice: Methods for Training and Inference using Low-Precision Numerics

Asit Mishra, Debbie Marr

PDF

Open Access

TL;DR

This paper introduces three methods for training and inference with low-precision numerics in deep learning, maintaining accuracy while reducing computational and memory costs, and presents an efficient hardware accelerator for these techniques.

Contribution

The paper proposes three novel schemes for low-precision training and inference that preserve accuracy and details an efficient hardware accelerator for implementation.

Findings

01

Low-precision numerics can be used without accuracy loss.

02

The proposed schemes improve computational efficiency.

03

An accelerator design optimizes low-precision deep learning.

Abstract

Today's high performance deep learning architectures involve large models with numerous parameters. Low precision numerics has emerged as a popular technique to reduce both the compute and memory requirements of these large models. However, lowering precision often leads to accuracy degradation. We describe three schemes whereby one can both train and do efficient inference using low precision numerics without hurting accuracy. Finally, we describe an efficient hardware accelerator that can take advantage of the proposed low precision numerics.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDistributed and Parallel Computing Systems