An Empirical Study on the Robustness of Massively Multilingual Neural   Machine Translation

Supryadi; Leiyu Pan; Deyi Xiong

arXiv:2405.07673·cs.CL·May 14, 2024

An Empirical Study on the Robustness of Massively Multilingual Neural Machine Translation

Supryadi, Leiyu Pan, Deyi Xiong

PDF

Open Access 1 Repo

TL;DR

This paper empirically examines the robustness of massively multilingual neural machine translation for Indonesian-Chinese, introducing a new benchmark dataset to evaluate translation quality under various natural noise conditions.

Contribution

It presents a novel robustness evaluation benchmark dataset for Indonesian-Chinese translation and analyzes the impact of noise and model size on translation robustness.

Findings

01

Correlation between error types and noise presence

02

Model size influences robustness and error patterns

03

Automatic evaluation metrics correlate with human judgments

Abstract

Massively multilingual neural machine translation (MMNMT) has been proven to enhance the translation quality of low-resource languages. In this paper, we empirically investigate the translation robustness of Indonesian-Chinese translation in the face of various naturally occurring noise. To assess this, we create a robustness evaluation benchmark dataset for Indonesian-Chinese translation. This dataset is automatically translated into Chinese using four NLLB-200 models of different sizes. We conduct both automatic and human evaluations. Our in-depth analysis reveal the correlations between translation error types and the types of noise present, how these correlations change across different model sizes, and the relationships between automatic evaluation indicators and human evaluation indicators. The dataset is publicly available at https://github.com/tjunlp-lab/ID-ZH-MTRobustEval.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

tjunlp-lab/id-zh-mtrobusteval
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques