Fluctuation-dissipation relations for stochastic gradient descent

Sho Yaida

arXiv:1810.00004·stat.ML·December 24, 2018·23 cites

Fluctuation-dissipation relations for stochastic gradient descent

Sho Yaida

PDF

Open Access 2 Repos

TL;DR

This paper derives fluctuation-dissipation relations for stochastic gradient descent that connect measurable quantities to hyperparameters, enabling adaptive training and insights into the loss landscape, with empirical validation.

Contribution

It introduces exact fluctuation-dissipation relations for SGD's stationary states, facilitating adaptive training and landscape analysis.

Findings

01

Relations hold exactly for any stationary state.

02

Can be used to adaptively set training schedules.

03

Efficiently extract Hessian and landscape information.

Abstract

The notion of the stationary equilibrium ensemble has played a central role in statistical mechanics. In machine learning as well, training serves as generalized equilibration that drives the probability distribution of model parameters toward stationarity. Here, we derive stationary fluctuation-dissipation relations that link measurable quantities and hyperparameters in the stochastic gradient descent algorithm. These relations hold exactly for any stationary state and can in particular be used to adaptively set training schedule. We can further use the relations to efficiently extract information pertaining to a loss-function landscape such as the magnitudes of its Hessian and anharmonicity. Our claims are empirically verified.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Markov Chains and Monte Carlo Methods · Gaussian Processes and Bayesian Inference