The Diffusion-Attention Connection

Julio Candanedo

arXiv:2604.09560·cs.LG·April 14, 2026

The Diffusion-Attention Connection

Julio Candanedo

PDF

TL;DR

This paper unifies transformers, diffusion-maps, and magnetic Laplacians into a single Markov geometric framework derived from query-scores, revealing their interconnectedness and dynamic regimes.

Contribution

It introduces a QK bidivergence that generalizes attention, diffusion-maps, and magnetic diffusion within a unified Markov geometry framework.

Findings

01

Unified view of attention and diffusion processes via QK bidivergence

02

Connections established between equilibrium, steady-state, and driven dynamics

03

Framework enables new insights into the geometric structure of neural models

Abstract

Transformers, diffusion-maps, and magnetic Laplacians are usually treated as separate tools; we show they are all different regimes of a single Markov geometry built from pre-softmax query-scores. We define a QK "bidivergence" whose exponentiated and normalized forms yield attention, diffusion-maps, and magnetic diffusion. And use product of experts and Schr\"odinger-bridges to connect and organize them into equilibrium, nonequilibrium steady-state, and driven dynamics.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.