Predicting the Emergence of Induction Heads in Language Model Pretraining

Tatsuya Aoyama; Ethan Gotlieb Wilcox; Nathan Schneider

arXiv:2511.16893·cs.CL·February 10, 2026

Predicting the Emergence of Induction Heads in Language Model Pretraining

Tatsuya Aoyama, Ethan Gotlieb Wilcox, Nathan Schneider

PDF

Open Access

TL;DR

This paper investigates how induction heads emerge in language models, revealing that their formation depends on training data statistics like bigram repetition and context size, regardless of model size.

Contribution

It provides a predictive equation for induction head emergence and analyzes how data properties influence their formation in language models.

Findings

01

Emergence point predicted by batch size and context size equation

02

Bigram repetition frequency and reliability strongly influence IH formation

03

Local dependency with high bigram repetition suffices for IH emergence

Abstract

Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in the context of language modeling, remains wanting. In this study, we investigate the relationship between statistical properties of the training data and IH formation in both natural and synthetic training data settings. We show that: (1) A simple equation combining batch size and context size predicts the point at which IHs form and that this emergence point is agnostic to model size; (2) Surface bigram repetition frequency and reliability strongly affect the formation of IHs, and we find an effective Pareto frontier in terms of these two values; (3) local dependency with high bigram repetition frequency and reliability is sufficient for IH formation, but when…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsLanguage and cultural evolution · Topic Modeling · Natural Language Processing Techniques