Towards Understanding Grokking: An Effective Theory of Representation   Learning

Ziming Liu; Ouail Kitouni; Niklas Nolte; Eric J. Michaud; Max Tegmark,; Mike Williams

arXiv:2205.10343·cs.LG·October 17, 2022·24 cites

Towards Understanding Grokking: An Effective Theory of Representation Learning

Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud, Max Tegmark,, Mike Williams

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper investigates the grokking phenomenon in deep learning, providing a theoretical framework and phase diagrams to explain how models generalize after overfitting, with insights into representation learning and phase transitions.

Contribution

It introduces an effective theory and phase diagram analysis to understand grokking, revealing the conditions for representation learning and the phases of learning in neural networks.

Findings

01

Grokking arises from structured representations predicted by the effective theory.

02

Four learning phases identified: comprehension, grokking, memorization, and confusion.

03

Representation learning occurs only within a specific 'Goldilocks zone' between memorization and confusion.

Abstract

We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set. We present both a microscopic analysis anchored by an effective theory and a macroscopic analysis of phase diagrams describing learning performance across hyperparameters. We find that generalization originates from structured representations whose training dynamics and dependence on training set size can be predicted by our effective theory in a toy setting. We observe empirically the presence of four learning phases: comprehension, grokking, memorization, and confusion. We find representation learning to occur only in a "Goldilocks zone" (including comprehension and grokking) between memorization and confusion. We find on transformers the grokking phase stays closer to the memorization phase (compared to the comprehension phase), leading to delayed generalization. The…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ejmichaud/grokking-squared
pytorchOfficial

Videos

Towards Understanding Grokking: An Effective Theory of Representation Learning· slideslive

Taxonomy

TopicsNeural dynamics and brain function · Neural Networks and Applications · Statistical Mechanics and Entropy