Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

Shaoke Fang; Ziang Li; Wenfei Wu; Jiatong Ji; Qingsong Liu; Ruizhi Pu

arXiv:2605.18825·cs.LG·May 20, 2026

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

Shaoke Fang, Ziang Li, Wenfei Wu, Jiatong Ji, Qingsong Liu, Ruizhi Pu

PDF

TL;DR

SAECache introduces a semantic-aware, adaptive eviction policy for LLM prefix caches, significantly improving reuse efficiency by learning token importance and workload characteristics online.

Contribution

It proposes a novel multi-queue, semantic-aware eviction policy with online learning, outperforming existing methods in LLM prefix cache management.

Findings

01

SAECache achieves 1.4x-2.7x TTFT improvement over baselines.

02

Fixed-parameter policies can degrade by up to 2.7x under workload mismatch.

03

Adaptive learning eliminates manual tuning and adapts to workload variations.

Abstract

Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt prefixes to reduce expensive prefill computation. However, its benefit depends critically on the eviction policy as GPU memory is scarce, and existing policies such as LRU largely treat cached blocks uniformly. This view ignores a fundamental property of LLM prompts: not all tokens are equally worth caching. We show that different token types within a prompt, including system prompts, user queries, tool outputs, model responses, and chain-of-thought reasoning, exhibit up to 756x variation in reuse rates, yet no existing eviction policy exploits this signal. In this paper, we present SAECache (Semantic-Adaptive Eviction for prefix caches), a semantic-adaptive prefix cache eviction policy that addresses this gap through three innovations:…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.