Gradient Descent on Neural Networks Typically Occurs at the Edge of   Stability

Jeremy M. Cohen; Simran Kaur; Yuanzhi Li; J. Zico Kolter; Ameet; Talwalkar

arXiv:2103.00065·cs.LG·November 24, 2022·22 cites

Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, Ameet, Talwalkar

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper empirically shows that neural network training with gradient descent often occurs at the Edge of Stability, where the Hessian's maximum eigenvalue hovers just above a critical value, challenging common optimization assumptions.

Contribution

It introduces the concept of the Edge of Stability in neural network training and provides empirical evidence of this regime's prevalence and characteristics.

Findings

01

Hessian eigenvalue hovers just above 2 / step size

02

Training loss exhibits non-monotonic short-term behavior

03

Long-term training loss decreases despite instability signs

Abstract

We empirically demonstrate that full-batch gradient descent on neural network training objectives typically operates in a regime we call the Edge of Stability. In this regime, the maximum eigenvalue of the training loss Hessian hovers just above the numerical value $2/ (step size)$ , and the training loss behaves non-monotonically over short timescales, yet consistently decreases over long timescales. Since this behavior is inconsistent with several widespread presumptions in the field of optimization, our findings raise questions as to whether these presumptions are relevant to neural network training. We hope that our findings will inspire future efforts aimed at rigorously understanding optimization at the Edge of Stability. Code is available at https://github.com/locuslab/edge-of-stability.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

locuslab/edge-of-stability
pytorchOfficial

Videos

Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability· slideslive

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Neural Networks and Applications · Machine Learning and Algorithms