Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

Maya Bechler-Speicher; Gilad Yehudai; Gil Harari; Clayton Sanford; Amir Globerson; Joan Bruna

arXiv:2605.22471·cs.LG·May 22, 2026

Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna

PDF

TL;DR

This paper investigates how different graph tokenizations affect transformer expressivity, revealing fundamental trade-offs and limitations in recovering structural information across tokenization types.

Contribution

It provides theoretical analysis of spectral, random-walk, and adjacency tokenizations, establishing depth regimes and impossibility results for transforming between them.

Findings

01

Random-walk tokenization is lossy for any walk length.

02

Spectral tokenization is lossless but ill-conditioned for local tasks.

03

Combining tokenizations improves structural signal extraction.

Abstract

Transformers have become a central architecture for graph learning, but their application to graphs requires first choosing a tokenization: a graph-to-token map that determines which structural information is exposed at the input. In this work, we show that this choice is a fundamental component of transformer expressivity. We examine three tokenizations that serve as building blocks for many existing graph tokenizations: spectral, random-walk, and adjacency tokenizations. We prove that different tokenizations induce distinct depth regimes: the same graph computation may be realizable by a shallow transformer under one tokenization, while requiring substantially larger depth under another. For example, we prove that random-walk tokenization is lossy for any walk length, making it impossible in general to recover the graph from it, and that while spectral tokenization is lossless, it is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.