The Historical Significance of Textual Distances

Ted Underwood

arXiv:1807.00181·cs.CL·July 3, 2018·1 cites

The Historical Significance of Textual Distances

Ted Underwood

PDF

Open Access 1 Repo

TL;DR

This paper investigates whether textual similarity measures accurately reflect cultural proximity by empirically comparing traditional and supervised learning-based methods in English fiction genres.

Contribution

It introduces new supervised learning strategies for measuring textual similarity anchored in social context, and compares them to existing methods.

Findings

01

Supervised learning methods better align with social measures.

02

Traditional cosine and topic vector similarities have limitations.

03

Empirical evidence supports the social relevance of new measures.

Abstract

Measuring similarity is a basic task in information retrieval, and now often a building-block for more complex arguments about cultural change. But do measures of textual similarity and distance really correspond to evidence about cultural proximity and differentiation? To explore that question empirically, this paper compares textual and social measures of the similarities between genres of English-language fiction. Existing measures of textual similarity (cosine similarity on tf-idf vectors or topic vectors) are also compared to new strategies that use supervised learning to anchor textual measurement in a social context.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

tedunderwood/genredistance
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsComputational and Text Analysis Methods · Data Analysis with R · Advanced Text Analysis Techniques