Evaluating historical text normalization systems: How well do they   generalize?

Alexander Robertson; Sharon Goldwater

arXiv:1804.02545·cs.CL·April 16, 2018

Evaluating historical text normalization systems: How well do they generalize?

Alexander Robertson, Sharon Goldwater

PDF

TL;DR

This paper critically examines how historical text normalization systems are evaluated, revealing that neural models generalize well but may not outperform naive baselines in downstream tasks, emphasizing the need for more rigorous evaluation practices.

Contribution

It identifies evaluation issues in historical text normalization and demonstrates the importance of comprehensive testing, including intrinsic and extrinsic measures.

Findings

01

Neural models generalize well to unseen words across five languages.

02

Neural models do not outperform naive baselines in downstream POS tagging.

03

Rigorous evaluation practices are necessary for assessing normalization systems.

Abstract

We highlight several issues in the evaluation of historical text normalization systems that make it hard to tell how well these systems would actually work in practice---i.e., for new datasets or languages; in comparison to more na\"ive systems; or as a preprocessing step for downstream NLP tools. We illustrate these issues and exemplify our proposed evaluation practices by comparing two neural models against a na\"ive baseline system. We show that the neural models generalize well to unseen words in tests on five languages; nevertheless, they provide no clear benefit over the na\"ive baseline for downstream POS tagging of an English historical collection. We conclude that future work should include more rigorous evaluation, including both intrinsic and extrinsic measures where possible.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.