On the Acceleration of Deep Neural Network Inference using Quantized   Compressed Sensing

Meshia C\'edric Oveneke

arXiv:2108.10101·cs.LG·August 24, 2021

On the Acceleration of Deep Neural Network Inference using Quantized Compressed Sensing

Meshia C\'edric Oveneke

PDF

Open Access

TL;DR

This paper introduces a novel binary quantization method based on quantized compressed sensing to accelerate DNN inference on resource-limited devices, aiming to reduce accuracy loss while improving speed and memory efficiency.

Contribution

It proposes a new binary quantization function leveraging quantized compressed sensing, theoretically reducing quantization error compared to standard methods.

Findings

01

The proposed QCS-based quantization reduces quantization error.

02

The method maintains the benefits of standard quantization techniques.

03

Theoretical analysis supports improved accuracy preservation.

Abstract

Accelerating deep neural network (DNN) inference on resource-limited devices is one of the most important barriers to ensuring a wider and more inclusive adoption. To alleviate this, DNN binary quantization for faster convolution and memory savings is one of the most promising strategies despite its serious drop in accuracy. The present paper therefore proposes a novel binary quantization function based on quantized compressed sensing (QCS). Theoretical arguments conjecture that our proposal preserves the practical benefits of standard methods, while reducing the quantization error and the resulting drop in accuracy.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSparse and Compressive Sensing Techniques · Advanced Image Processing Techniques · Image Processing Techniques and Applications

MethodsConvolution