# Wikipedia Cultural Diversity Dataset: A Complete Cartography for 300   Language Editions

**Authors:** Marc Miquel-Rib\'e, David Laniado

arXiv: 1901.07999 · 2019-06-11

## TL;DR

This paper introduces the Wikipedia Cultural Diversity dataset, classifying articles across 300 language editions to analyze cultural representation and support cross-cultural research in digital humanities.

## Contribution

It provides a comprehensive dataset with classification methodology and features for cultural context articles across multiple Wikipedia language editions.

## Key findings

- Dataset covers 300 language editions.
- Methodology for classifying cultural articles.
- Potential applications in content gap analysis.

## Abstract

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and entities related to the language and to the territories where it is spoken. We describe the methodology we employed to classify articles, and the rich set of features that we defined to feed the classifier, and that are released as part of the dataset. We present several purposes for which we envision the use of this dataset, including detecting, measuring and countering content gaps in the Wikipedia project, and encouraging cross-cultural research in the field of digital humanities.

---
Source: https://tomesphere.com/paper/1901.07999