Mini-batch Serialization: CNN Training with Inter-layer Data Reuse

Sangkug Lym; Armand Behroozi; Wei Wen; Ge Li; Yongkee Kwon; Mattan; Erez

arXiv:1810.00307·cs.LG·May 7, 2019·1 cites

Mini-batch Serialization: CNN Training with Inter-layer Data Reuse

Sangkug Lym, Armand Behroozi, Wei Wen, Ge Li, Yongkee Kwon, Mattan, Erez

PDF

Open Access 1 Repo

TL;DR

This paper proposes Mini-batch Serialization (MBS), a novel approach that reorganizes CNN training to significantly reduce memory traffic and improve efficiency by better utilizing on-chip buffers and inter-layer data reuse.

Contribution

It introduces the MBS training method and the WaveCore accelerator, achieving substantial reductions in memory traffic and energy consumption while boosting performance.

Findings

01

DRAM traffic reduced by 75%

02

Performance improved by 53%

03

System energy saved by 26%

Abstract

Training convolutional neural networks (CNNs) requires intense computations and high memory bandwidth. We find that bandwidth today is over-provisioned because most memory accesses in CNN training can be eliminated by rearranging computation to better utilize on-chip buffers and avoid traffic resulting from large per-layer memory footprints. We introduce the MBS CNN training approach that significantly reduces memory traffic by partially serializing mini-batch processing across groups of layers. This optimizes reuse within on-chip buffers and balances both intra-layer and inter-layer reuse. We also introduce the WaveCore CNN training accelerator that effectively trains CNNs in the MBS approach with high functional-unit utilization. Combined, WaveCore and MBS reduce DRAM traffic by 75%, improve performance by 53%, and save 26% system energy for modern deep CNN training compared to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

https://bitbucket.org/lph_tools/mini-batch-serialization
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Neural Network Applications · Ferroelectric and Negative Capacitance Devices · Stochastic Gradient Optimization Techniques