Scalable K-FAC Training for Deep Neural Networks with Distributed   Preconditioning

Lin Zhang; Shaohuai Shi; Wei Wang; Bo Li

arXiv:2206.15143·cs.LG·July 1, 2022

Scalable K-FAC Training for Deep Neural Networks with Distributed Preconditioning

Lin Zhang, Shaohuai Shi, Wei Wang, Bo Li

PDF

Open Access 1 Repo

TL;DR

This paper introduces DP-KFAC, a distributed preconditioning method for second-order DNN training that reduces computation, communication, and memory costs while maintaining convergence, demonstrated on a 64-GPU cluster.

Contribution

DP-KFAC distributes Kronecker factor construction across workers, significantly improving efficiency over existing D-KFAC algorithms.

Findings

01

Reduces computation overhead by up to 1.65x.

02

Decreases communication costs by up to 3.15x.

03

Lowers memory footprint in second-order updates.

Abstract

The second-order optimization methods, notably the D-KFAC (Distributed Kronecker Factored Approximate Curvature) algorithms, have gained traction on accelerating deep neural network (DNN) training on GPU clusters. However, existing D-KFAC algorithms require to compute and communicate a large volume of second-order information, i.e., Kronecker factors (KFs), before preconditioning gradients, resulting in large computation and communication overheads as well as a high memory footprint. In this paper, we propose DP-KFAC, a novel distributed preconditioning scheme that distributes the KF constructing tasks at different DNN layers to different workers. DP-KFAC not only retains the convergence property of the existing D-KFAC algorithms but also enables three benefits: reduced computation overhead in constructing KFs, no communication of KFs, and low memory footprint. Extensive experiments on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

lzhangbv/kfac_pytorch
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Advanced Neural Network Applications · Machine Learning and ELM