Spherical Motion Dynamics: Learning Dynamics of Neural Network with   Normalization, Weight Decay, and SGD

Ruosi Wan; Zhanxing Zhu; Xiangyu Zhang; Jian Sun

arXiv:2006.08419·stat.ML·November 30, 2020·6 cites

Spherical Motion Dynamics: Learning Dynamics of Neural Network with Normalization, Weight Decay, and SGD

Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian Sun

PDF

Open Access

TL;DR

This paper investigates the learning dynamics of neural networks with normalization, weight decay, and SGD, revealing conditions for equilibrium and introducing angular update as a new measure, supported by theoretical proofs and experiments.

Contribution

It introduces assumptions leading to equilibrium in SMD, proves linear convergence of weight norm and angular update, and validates findings on large-scale vision tasks.

Findings

01

Weight norm converges at a linear rate under certain assumptions.

02

Angular update can be used as a measure and converges linearly.

03

Theoretical results align well with empirical observations on ImageNet and MSCOCO.

Abstract

In this work, we comprehensively reveal the learning dynamics of neural network with normalization, weight decay (WD), and SGD (with momentum), named as Spherical Motion Dynamics (SMD). Most related works study SMD by focusing on "effective learning rate" in "equilibrium" condition, where weight norm remains unchanged. However, their discussions on why equilibrium condition can be reached in SMD is either absent or less convincing. Our work investigates SMD by directly exploring the cause of equilibrium condition. Specifically, 1) we introduce the assumptions that can lead to equilibrium condition in SMD, and prove that weight norm can converge at linear rate with given assumptions; 2) we propose "angular update" as a substitute for effective learning rate to measure the evolving of neural network in SMD, and prove angular update can also converge to its theoretical value at linear…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Model Reduction and Neural Networks · Advanced Neural Network Applications

MethodsStochastic Gradient Descent · Weight Decay · Batch Normalization