Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension

Michael Crawshaw; Mingrui Liu

arXiv:2602.12471·cs.LG·February 16, 2026

Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension

Michael Crawshaw, Mingrui Liu

PDF

Open Access

TL;DR

This paper provides a precise analysis of gradient descent for logistic regression in two dimensions with large step sizes, establishing tight bounds on convergence time and loss reduction, especially during the transition from unstable to stable phases.

Contribution

It offers a tighter, dimension-specific analysis of GD dynamics for logistic regression, including matching lower bounds, improving understanding of convergence behavior with large step sizes.

Findings

01

GD with large step size achieves loss below O(1/(ta T))

02

Transition time from unstable to stable loss is tightly bounded

03

Analysis is tight up to logarithmic factors

Abstract

We consider the optimization problem of minimizing the logistic loss with gradient descent to train a linear model for binary classification with separable data. With a budget of $T$ iterations, it was recently shown that an accelerated $1/ T^{2}$ rate is possible by choosing a large step size $η = Θ (γ^{2} T)$ (where $γ$ is the dataset's margin) despite the resulting non-monotonicity of the loss. In this paper, we provide a tighter analysis of gradient descent for this problem when the data is two-dimensional: we show that GD with a sufficiently large learning rate $η$ finds a point with loss smaller than $O (1/ (η T))$ , as long as $T \geq Ω (n / γ + 1/ γ^{2})$ , where $n$ is the dataset size. Our improved rate comes from a tighter bound on the time $τ$ that it takes for GD to transition from unstable (non-monotonic loss) to stable (monotonic loss),…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Sparse and Compressive Sensing Techniques · Privacy-Preserving Technologies in Data