Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution

Rong Feng; Suman Saha

arXiv:2511.19130·cs.SE·November 25, 2025

Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution

Rong Feng, Suman Saha

PDF

Open Access

TL;DR

This paper evaluates whether fine-tuned large language models, enhanced with symbolic execution artifacts, can effectively deobfuscate and recover the semantics of obfuscated C programs, improving software analysis tasks.

Contribution

It systematically assesses the effectiveness of LLMs combined with symbolic execution artifacts in deobfuscating code, introducing a benchmark with diverse obfuscation techniques and training configurations.

Findings

01

GPT-4.1-mini outperforms other models in deobfuscation accuracy

02

Incorporating KLEE artifacts improves semantic fidelity and compilation success

03

Symbolic execution artifacts enhance LLMs' ability to recover program semantics

Abstract

Obfuscation poses a persistent challenge for software engineering tasks such as program comprehension, maintenance, testing, and vulnerability detection. While compiler optimizations and third-party code often introduce transformations that obscure program intent, existing analysis tools and large language models (LLMs) struggle to recover the original semantics. In this work, we investigate whether LLMs, when fine-tuned with symbolic execution artifacts, can effectively deobfuscate programs and restore analyzability. We construct a benchmark by applying four widely studied transformations-control-flow flattening, opaque predicates, arithmetic encoding, and branch encoding-across diverse C programs from TUM Obfuscation Benchmarks, the LLVM test suite, and algorithmic repositories. We then compare three state-of-the-art LLMs under two training configurations: baseline fine-tuning on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSoftware Testing and Debugging Techniques · Software Engineering Research · Advanced Malware Detection Techniques