Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences

Neehal Tumma; Noel Loo; Daniela Rus

arXiv:2604.21100·cs.LG·April 24, 2026

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences

Neehal Tumma, Noel Loo, Daniela Rus

PDF

TL;DR

This paper introduces preconditioning to delta-rule recurrences, enhancing their performance in sequence modeling tasks by incorporating curvature information from online least squares theory.

Contribution

It develops a theoretical framework connecting linear attention and delta rule with preconditioning, and implements practical preconditioned variants of DeltaNet, GDN, and KDA.

Findings

01

Preconditioned delta-rule recurrences improve performance on synthetic recall benchmarks.

02

They also enhance language modeling at 340M and 1B scale.

Abstract

To address the increasing long-context compute limitations of softmax attention, several subquadratic recurrent operators have been developed. This work includes models such as Mamba-2, DeltaNet, Gated DeltaNet (GDN), and Kimi Delta Attention (KDA). As the space of recurrences grows, a parallel line of work has arisen to taxonomize them. One compelling view is the test-time regression (TTR) framework, which interprets recurrences as performing online least squares updates that learn a linear map from the keys to values. Existing delta-rule recurrences can be seen as first-order approximations to this objective, but notably ignore the curvature of the least-squares loss during optimization. In this work, we address this by introducing preconditioning to these recurrences. Starting from the theory of online least squares, we derive equivalences between linear attention and the delta rule…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.