Machine Translation Customization via Automatic Training Data Selection   from the Web

Thuy Vu; Alessandro Moschitti

arXiv:2102.10243·cs.CL·February 23, 2021

Machine Translation Customization via Automatic Training Data Selection from the Web

Thuy Vu, Alessandro Moschitti

PDF

1 Repo

TL;DR

This paper presents a method for customizing neural machine translation systems to specific domains by automatically selecting relevant training data from the Web using document classifiers, leading to improved performance with less data.

Contribution

The authors introduce a novel data selection approach using monolingual target data and document classifiers to enhance domain-specific neural machine translation.

Findings

01

Outperforms top systems on WMT-18 News translation benchmark

02

Uses less data and smaller models than state-of-the-art systems

03

Achieves better domain adaptation for machine translation

Abstract

Machine translation (MT) systems, especially when designed for an industrial setting, are trained with general parallel data derived from the Web. Thus, their style is typically driven by word/structure distribution coming from the average of many domains. In contrast, MT customers want translations to be specialized to their domain, for which they are typically able to provide text samples. We describe an approach for customizing MT systems on specific domains by selecting data similar to the target customer data to train neural translation models. We build document classifiers using monolingual target data, e.g., provided by the customers to select parallel training data from Web crawled data. Finally, we train MT models on our automatically selected data, obtaining a system specialized to the target domain. We tested our approach on the benchmark from WMT-18 Translation Task for News…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

awslabs/sockeye
mxnetOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.