Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and   Luck

Benjamin L. Edelman; Surbhi Goel; Sham Kakade; Eran Malach; Cyril; Zhang

arXiv:2309.03800·cs.LG·October 31, 2023

Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and Luck

Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, Cyril, Zhang

PDF

Open Access

TL;DR

This paper explores how neural network architecture choices, like width and initialization, influence resource tradeoffs in feature learning, demonstrating that wider, sparsely-initialized models improve sample efficiency and can outperform traditional methods on benchmarks.

Contribution

It introduces a theoretical and experimental framework linking network width and initialization to resource tradeoffs and sample efficiency in feature learning, especially in sparse parity tasks.

Findings

01

Wider networks improve sample efficiency in sparse parity learning.

02

Sparse initialization enhances the probability of finding lottery ticket neurons.

03

Wide, sparsely-initialized models can outperform tuned random forests on benchmarks.

Abstract

In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learning in the presence of computational-statistical gaps. We begin by considering offline sparse parity learning, a supervised classification problem which admits a statistical query lower bound for gradient-based training of a multilayer perceptron. This lower bound can be interpreted as a multi-resource tradeoff frontier: successful learning can only occur if one is sufficiently rich (large model), knowledgeable (large dataset), patient (many training iterations), or lucky (many random guesses). We show, theoretically and experimentally, that sparse initialization and increasing network width yield significant improvements in sample efficiency in this setting. Here, width…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Neural Networks and Applications · Machine Learning and Algorithms