The Convergence Rate of Neural Networks for Learned Functions of   Different Frequencies

Ronen Basri; David Jacobs; Yoni Kasten; Shira Kritchman

arXiv:1906.00425·cs.LG·December 3, 2019·90 cites

The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies

Ronen Basri, David Jacobs, Yoni Kasten, Shira Kritchman

PDF

Open Access 1 Repo

TL;DR

This paper investigates how the frequency of a function affects the learning speed of neural networks, deriving theoretical predictions that match empirical results for both shallow and deep models.

Contribution

It introduces a theoretical framework linking function frequency to neural network convergence rates, including the impact of bias terms on low-frequency function learning.

Findings

01

Bias terms are crucial for learning odd-frequency functions.

02

Theoretical predictions align with empirical learning times.

03

Low-frequency functions are learned faster than high-frequency ones.

Abstract

We study the relationship between the frequency of a function and the speed at which a neural network learns it. We build on recent results that show that the dynamics of overparameterized neural networks trained with gradient descent can be well approximated by a linear system. When normalized training data is uniformly distributed on a hypersphere, the eigenfunctions of this linear system are spherical harmonic functions. We derive the corresponding eigenvalues for each frequency after introducing a bias term in the model. This bias term had been omitted from the linear network model without significantly affecting previous theoretical results. However, we show theoretically and experimentally that a shallow neural network without bias cannot represent or learn simple, low frequency functions with odd frequencies. Our results lead to specific predictions of the time it will take a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ykasten/Convergence-Rate-NN-Different-Frequencies
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications · Model Reduction and Neural Networks · Stochastic Gradient Optimization Techniques

MethodsSPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings