SGD Through the Lens of Kolmogorov Complexity

Gregory Schwartzman

arXiv:2111.05478·cs.LG·May 17, 2022

SGD Through the Lens of Kolmogorov Complexity

Gregory Schwartzman

PDF

Open Access

TL;DR

This paper proves that stochastic gradient descent (SGD) can achieve near-perfect classification accuracy under mild assumptions, including low model complexity and local progress, providing the first convergence guarantee for general, underparameterized models.

Contribution

It introduces a novel convergence analysis of SGD based on Kolmogorov complexity, applicable to a wide range of models without specific architectural constraints.

Findings

01

SGD achieves (1-ε) accuracy under mild assumptions.

02

First convergence guarantee for underparameterized models.

03

Model-agnostic analysis using entropy compression.

Abstract

We prove that stochastic gradient descent (SGD) finds a solution that achieves $(1 - ϵ)$ classification accuracy on the entire dataset. We do so under two main assumptions: (1. Local progress) The model accuracy improves on average over batches. (2. Models compute simple functions) The function computed by the model is simple (has low Kolmogorov complexity). It is sufficient that these assumptions hold only for a tiny fraction of the epochs. Intuitively, the above means that intermittent local progress of SGD implies global progress. Assumption 2 trivially holds for underparameterized models, hence, our work gives the first convergence guarantee for general, underparameterized models. Furthermore, this is the first result which is completely model agnostic - we do not require the model to have any specific architecture or activation function, it may not even be a neural network.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Markov Chains and Monte Carlo Methods · Machine Learning and Algorithms

MethodsStochastic Gradient Descent