From Sublinear to Linear: Fast Convergence in Deep Networks via Locally Polyak-Lojasiewicz Regions

Agnideep Aich; Ashit Baran Aich; Bruce Wade

arXiv:2507.21429·stat.ML·January 13, 2026

From Sublinear to Linear: Fast Convergence in Deep Networks via Locally Polyak-Lojasiewicz Regions

Agnideep Aich, Ashit Baran Aich, Bruce Wade

PDF

TL;DR

This paper demonstrates that under certain local stability conditions, gradient descent on deep networks exhibits linear convergence within locally Polyak-Lojasiewicz regions, explaining fast training dynamics observed in practice.

Contribution

It introduces the concept of LPLRs under local NTK stability, establishing linear convergence of GD in deep networks and providing empirical evidence with modern architectures.

Findings

01

GD converges linearly within LPLRs under local NTK stability

02

Empirical evidence of PL-like scaling in CNN training

03

LPLR signatures persist in stochastic optimization settings

Abstract

Gradient descent (GD) on deep neural network loss landscapes is non-convex, yet often converges far faster in practice than classical guarantees suggest. Prior work shows that within locally quasi-convex regions (LQCRs), GD converges to stationary points at sublinear rates, leaving the commonly observed near-exponential training dynamics unexplained. We show that, under a mild local Neural Tangent Kernel (NTK) stability assumption, the loss satisfies a PL-type error bound within these regions, yielding a Locally Polyak-Lojasiewicz Region (LPLR) in which the squared gradient norm controls the suboptimality gap. For properly initialized finite-width networks, we show that under local NTK stability this PL-type mechanism holds around initialization and establish linear convergence of GD as long as the iterates remain within the resulting LPLR. Empirically, we observe PL-like scaling and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.