Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

Tengyuan Liang; Alexander Rakhlin

arXiv:1808.00387·math.ST·July 27, 2020

Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

Tengyuan Liang, Alexander Rakhlin

PDF

TL;DR

This paper investigates how kernel ridgeless regression can generalize well without explicit regularization, due to implicit regularization effects from data geometry, kernel curvature, and high dimensionality, supported by theoretical bounds and experiments.

Contribution

It identifies and explains the implicit regularization mechanism enabling kernel ridgeless regression to generalize, with theoretical bounds and empirical evidence.

Findings

01

Implicit regularization arises from data geometry and kernel properties.

02

Theoretical upper bounds on out-of-sample error are derived.

03

Experimental results on MNIST support the phenomenon.

Abstract

In the absence of explicit regularization, Kernel "Ridgeless" Regression with nonlinear kernels has the potential to fit the training data perfectly. It has been observed empirically, however, that such interpolated solutions can still generalize well on test data. We isolate a phenomenon of implicit regularization for minimum-norm interpolated solutions which is due to a combination of high dimensionality of the input data, curvature of the kernel function, and favorable geometric properties of the data such as an eigenvalue decay of the empirical covariance and kernel matrices. In addition to deriving a data-dependent upper bound on the out-of-sample error, we present experimental evidence suggesting that the phenomenon occurs in the MNIST dataset.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.