# Wikipedia Text Reuse: Within and Without

**Authors:** Milad Alshomary, Michael V\"olske, Tristan Licht, Henning Wachsmuth,, Benno Stein, Matthias Hagen, Martin Potthast

arXiv: 1812.09221 · 2018-12-24

## TL;DR

This paper presents a large-scale analysis of text reuse within and outside Wikipedia, using advanced detection technology to uncover millions of reuse cases, and provides a new corpus and pipeline for future research.

## Contribution

It introduces the first comprehensive corpus of Wikipedia text reuse cases and scales state-of-the-art detection methods to analyze reuse across the entire platform and web.

## Key findings

- Discovered 100 million reuse cases inside Wikipedia
- Identified 1.6 million reuse cases outside Wikipedia
- Enabled new tasks like template induction and quality correction

## Abstract

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy and paste, we employ state-of-the-art text reuse detection technology, scaling it for the first time to process the entire Wikipedia as part of a distributed retrieval pipeline. We further report on a pilot analysis of the 100 million reuse cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered. Text reuse inside Wikipedia gives rise to new tasks such as article template induction, fixing quality flaws due to inconsistencies arising from asynchronous editing of reused passages, or complementing Wikipedia's ontology. Text reuse outside Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia's influence on the web. To foster future research into these tasks, and for reproducibility's sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely available.

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/1812.09221/full.md

## Figures

5 figures with captions in the complete paper: https://tomesphere.com/paper/1812.09221/full.md

## References

20 references — full list in the complete paper: https://tomesphere.com/paper/1812.09221/full.md

---
Source: https://tomesphere.com/paper/1812.09221