Agnostic Learning of General ReLU Activation Using Gradient Descent

Pranjal Awasthi; Alex Tang; Aravindan Vijayaraghavan

arXiv:2208.02711·cs.LG·November 5, 2024

Agnostic Learning of General ReLU Activation Using Gradient Descent

Pranjal Awasthi, Alex Tang, Aravindan Vijayaraghavan

PDF

Open Access 1 Video

TL;DR

This paper analyzes the convergence of gradient descent for agnostically learning ReLU functions with non-zero bias under Gaussian distributions, providing guarantees on the learned function's error relative to the best possible ReLU.

Contribution

It extends previous analyses to include non-zero bias ReLUs, offering finite sample guarantees and broader distribution applicability.

Findings

01

Gradient descent converges to a near-optimal ReLU with high probability.

02

The analysis applies to a broader class of distributions beyond Gaussians.

03

Finite sample guarantees are established for the learning process.

Abstract

We provide a convergence analysis of gradient descent for the problem of agnostically learning a single ReLU function with moderate bias under Gaussian distributions. Unlike prior work that studies the setting of zero bias, we consider the more challenging scenario when the bias of the ReLU function is non-zero. Our main result establishes that starting from random initialization, in a polynomial number of iterations gradient descent outputs, with high probability, a ReLU function that achieves an error that is within a constant factor of the optimal error of the best ReLU function with moderate bias. We also provide finite sample guarantees, and these techniques generalize to a broader class of marginal distributions beyond Gaussians.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Agnostic Learning of General ReLU Activation Using Gradient Descent· slideslive

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Machine Learning and ELM · Machine Learning and Algorithms