Transformer Block Coupling and its Correlation with Generalization in   LLMs

Murdock Aubry; Haoming Meng; Anton Sugolov; Vardan Papyan

arXiv:2407.07810·cs.LG·March 6, 2025

Transformer Block Coupling and its Correlation with Generalization in LLMs

Murdock Aubry, Haoming Meng, Anton Sugolov, Vardan Papyan

PDF

Open Access

TL;DR

This paper investigates the internal dynamics of transformer blocks in LLMs, revealing a phenomenon called transformer block coupling that correlates positively with model performance and generalization, supported by empirical analysis and experiments.

Contribution

It introduces the concept of transformer block coupling, analyzes its emergence during training, and links it to improved generalization in both language and vision transformers.

Findings

01

Coupling correlates positively with model performance.

02

Coupling and linearity increase during training.

03

Coupling observed in both LLMs and Vision Transformers.

Abstract

Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential. In this work, we analyze the trajectories of token embeddings as they pass through transformer blocks, linearizing the system along these trajectories through their Jacobian matrices. By examining the relationships between these block Jacobians, we uncover the phenomenon of \textbf{transformer block coupling} in a multitude of LLMs, characterized by the coupling of their top singular vectors across tokens and depth. Our findings reveal that coupling \textit{positively correlates} with model performance, and that this relationship is stronger than with other hyperparameters such as parameter count, model depth, and embedding dimension. We further investigate how these properties emerge during training,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling