An Information-Theoretic Approach for Detecting Edits in AI-Generated   Text

Idan Kashtan; Alon Kipnis

arXiv:2308.12747·cs.IT·August 27, 2024

An Information-Theoretic Approach for Detecting Edits in AI-Generated Text

Idan Kashtan, Alon Kipnis

PDF

Open Access

TL;DR

This paper introduces an information-theoretic method to detect whether text is AI-generated or edited by humans, effectively identifying the origin of sentences and pinpointing edits within the text.

Contribution

It presents a novel, sensitive testing approach that combines multiple tests to determine text origin and detect edits, grounded in an information-theoretic framework.

Findings

01

Effective detection of AI-generated text and edits demonstrated through extensive real-data evaluations.

02

The method identifies rare, scattered non-null effects indicating edits or different origins.

03

Theoretical analysis highlights optimality conditions and raises new research questions.

Abstract

We propose a method to determine whether a given article was written entirely by a generative language model or perhaps contains edits by a different author, possibly a human. Our process involves multiple tests for the origin of individual sentences or other pieces of text and combining these tests using a method that is sensitive to rare alternatives, i.e., non-null effects are few and scattered across the text in unknown locations. Interestingly, this method also identifies pieces of text suspected to contain edits. We demonstrate the effectiveness of the method in detecting edits through extensive evaluations using real data and provide an information-theoretic analysis of the factors affecting its success. In particular, we discuss optimality properties under a theoretical framework for text editing saying that sentences are generated mainly by the language model, except perhaps…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling