The Impact of the Mini-batch Size on the Variance of Gradients in   Stochastic Gradient Descent

Xin Qian; Diego Klabjan

arXiv:2004.13146·math.OC·April 29, 2020·29 cites

The Impact of the Mini-batch Size on the Variance of Gradients in Stochastic Gradient Descent

Xin Qian, Diego Klabjan

PDF

Open Access

TL;DR

This paper analyzes how mini-batch size affects gradient variance in stochastic gradient descent, providing theoretical insights and empirical evidence that smaller batches lead to lower gradient variance and better training outcomes.

Contribution

It offers the first theoretical analysis of gradient variance dependence on mini-batch size in linear and deep linear networks, supported by empirical validation.

Findings

01

Gradient norm decreases with larger mini-batch size

02

Gradient variance is a decreasing function of batch size

03

Smaller batches tend to produce lower loss values

Abstract

The mini-batch stochastic gradient descent (SGD) algorithm is widely used in training machine learning models, in particular deep learning models. We study SGD dynamics under linear regression and two-layer linear networks, with an easy extension to deeper linear networks, by focusing on the variance of the gradients, which is the first study of this nature. In the linear regression case, we show that in each iteration the norm of the gradient is a decreasing function of the mini-batch size $b$ and thus the variance of the stochastic gradient estimator is a decreasing function of $b$ . For deep neural networks with $L_{2}$ loss we show that the variance of the gradient is a polynomial in $1/ b$ . The results back the important intuition that smaller batch sizes yield lower loss function values which is a common believe among the researchers. The proof techniques exhibit a relationship…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Advanced Neural Network Applications · Machine Learning and ELM

MethodsLinear Regression · Stochastic Gradient Descent