Remember to Forget: Gated Adaptive Positional Encoding

Riccardo Ali; Alessio Borgi; Christopher Irwin; Mario Severino; Pietro Li\`o

arXiv:2605.10414·cs.LG·May 12, 2026

Remember to Forget: Gated Adaptive Positional Encoding

Riccardo Ali, Alessio Borgi, Christopher Irwin, Mario Severino, Pietro Li\`o

PDF

TL;DR

GAPE is a novel positional encoding method that enhances long-context robustness in language models by introducing content-aware, gate-based attention modulation while maintaining rotary geometry.

Contribution

It proposes GAPE, a drop-in augmentation for rotary positional encoding that improves long-range attention stability without sacrificing local resolution.

Findings

01

GAPE yields sharper attention compared to rotary baselines.

02

GAPE improves robustness in synthetic retrieval and long-context benchmarks.

03

GAPE can be implemented within standard scaled dot-product attention.

Abstract

Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during training, rotary phases can enter out-of-distribution regimes, leading to spurious long-range alignments, diffuse attention, and degraded retrieval. Existing remedies only partially address these failures, as they often trade local positional resolution for long-context stability. We propose GAPE (Gated Adaptive Positional Encoding), a drop-in augmentation for positional encodings that introduces a content-aware bias directly into the attention logits while preserving the rotary geometry. GAPE decouples distance-based suppression from token importance through a query-dependent gate that contracts irrelevant context and a key-dependent gate that preserves salient distant tokens. We prove that protected tokens remain accessible, while the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.