Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning

Adam \v{S}torek; Mukur Gupta; Samira Hajizadeh; Prashast Srivastava; Suman Jana

arXiv:2505.13353·cs.CL·April 21, 2026

Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning

Adam \v{S}torek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava, Suman Jana

PDF

TL;DR

This paper investigates whether large language models truly understand code semantics in long contexts or rely on pattern matching, revealing significant semantic recall degradation and positional effects.

Contribution

It introduces semantic recall sensitivity, a novel measurement method, and a new task SemTrace to evaluate true semantic understanding in LLMs.

Findings

01

Frontier models excel at lexical recall but struggle with semantic recall in long contexts.

02

Models heavily depend on pattern matching shortcuts, especially when relevant code is centrally positioned.

03

Semantic recall sensitivity reveals severe positional accuracy drops, indicating underestimated semantic understanding failures.

Abstract

Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long code context or rely on pattern matching shortcuts remains unclear. We distinguish between lexical recall (retrieving code verbatim) and semantic recall (understanding operational semantics). Evaluating 10 state-of-the-art LLMs, we find that while frontier models achieve near-perfect, position-independent lexical recall, semantic recall degrades severely when code is centrally positioned in long contexts. We introduce semantic recall sensitivity to measure whether tasks require understanding of code's operational semantics vs. permit pattern matching shortcuts. Through a novel counterfactual measurement method, we show that models rely heavily on pattern matching shortcuts to solve existing code understanding benchmarks. We propose a new…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.