Non-Euclidean Gradient Descent Operates at the Edge of Stability

Rustem Islamov; Michael Crawshaw; Jeremy Cohen; Robert Gower

arXiv:2603.05002·cs.LG·March 6, 2026

Non-Euclidean Gradient Descent Operates at the Edge of Stability

Rustem Islamov, Michael Crawshaw, Jeremy Cohen, Robert Gower

PDF

Open Access

TL;DR

This paper extends the understanding of the Edge of Stability phenomenon in gradient descent to non-Euclidean norms, providing a unified, geometry-aware spectral measure that explains training dynamics across various optimizers.

Contribution

It introduces a generalized sharpness measure applicable to arbitrary norms, extending the EoS analysis beyond Euclidean settings and encompassing multiple optimization methods.

Findings

01

Non-Euclidean GD exhibits EoS behavior similar to Euclidean GD.

02

The generalized sharpness measure captures optimizer dynamics across different geometries.

03

Experiments confirm progressive sharpening and oscillations around the stability threshold.

Abstract

The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian converges to $2/ η$ during training with gradient descent (GD) with a step-size $η$ . Despite (apparently) violating classical smoothness assumptions, EoS has been widely observed in deep learning, but its theoretical foundations remain incomplete. We provide an interpretation of EoS through the lens of Directional Smoothness Mishkin et al. [2024]. This interpretation naturally extends to non-Euclidean norms, which we use to define generalized sharpness under an arbitrary norm. Our generalized sharpness measure includes previously studied vanilla GD and preconditioned GD as special cases, as well as methods for which EoS has not been studied, such as $ℓ_{\infty}$ -descent, Block CD, Spectral GD, and Muon without momentum. Through experiments on neural networks, we show that…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Machine Learning in Materials Science · Adversarial Robustness in Machine Learning