Feature Learning Beyond the Edge of Stability

D\'avid Terj\'ek

arXiv:2502.13110·cs.LG·May 20, 2025

Feature Learning Beyond the Edge of Stability

D\'avid Terj\'ek

PDF

Open Access

TL;DR

This paper introduces a new neural network training method that leverages a polynomial width pattern and depthwise gradient scaling to enhance feature learning and stability beyond traditional limits, supported by theoretical formulas and empirical results.

Contribution

It presents a novel homogeneous multilayer perceptron parameterization with a specific width pattern and gradient scaling scheme, enabling stable training beyond the edge of stability.

Findings

01

Improved feature learning demonstrated empirically.

02

Gradient scaling scheme enables stable training beyond stability edge.

03

Theoretical formulas connect sharpness and feature quality.

Abstract

We propose a homogeneous multilayer perceptron parameterization with polynomial hidden layer width pattern and analyze its training dynamics under stochastic gradient descent with depthwise gradient scaling in a general supervised learning scenario. We obtain formulas for the first three Taylor coefficients of the minibatch loss during training that illuminate the connection between sharpness and feature learning, providing in particular a soft rank variant that quantifies the quality of learned hidden layer features. Based on our theory, we design a gradient scaling scheme that in tandem with a quadratic width pattern enables training beyond the edge of stability without loss explosions or numerical errors, resulting in improved feature learning and implicit sharpness regularization as demonstrated empirically.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques