When Coding Style Survives Compilation: De-anonymizing Programmers from   Executable Binaries

Aylin Caliskan; Fabian Yamaguchi; Edwin Dauber; Richard Harang; Konrad; Rieck; Rachel Greenstadt; and Arvind Narayanan

arXiv:1512.08546·cs.CR·December 19, 2017

When Coding Style Survives Compilation: De-anonymizing Programmers from Executable Binaries

Aylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard Harang, Konrad, Rieck, Rachel Greenstadt, and Arvind Narayanan

PDF

3 Repos

TL;DR

This paper demonstrates that machine learning techniques can effectively de-anonymize programmers from executable binaries, even after obfuscation and stripping, raising privacy concerns for developers.

Contribution

It introduces a novel approach combining decompilation and stylistic features for binary authorship attribution, achieving high accuracy and robustness against obfuscation.

Findings

01

Achieved up to 96% attribution accuracy on Google Code Jam data.

02

Robust to obfuscation, compiler optimizations, and stripped binaries.

03

Effective on real-world code from GitHub and hacker forums.

Abstract

The ability to identify authors of computer programs based on their coding style is a direct threat to the privacy and anonymity of programmers. While recent work found that source code can be attributed to authors with high accuracy, attribution of executable binaries appears to be much more difficult. Many distinguishing features present in source code, e.g. variable names, are removed in the compilation process, and compiler optimization may alter the structure of a program, further obscuring features that are known to be useful in determining authorship. We examine programmer de-anonymization from the standpoint of machine learning, using a novel set of features that include ones obtained by decompiling the executable binary to source code. We adapt a powerful set of techniques from the domain of source code authorship attribution along with stylistic representations embedded in…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.