Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation

Diptesh Kanojia; Marina Fomicheva; Tharindu Ranasinghe; Fr\'ed\'eric; Blain; Constantin Or\u{a}san; Lucia Specia

arXiv:2109.10859·cs.CL·September 23, 2021

Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation

Diptesh Kanojia, Marina Fomicheva, Tharindu Ranasinghe, Fr\'ed\'eric, Blain, Constantin Or\u{a}san, Lucia Specia

PDF

Open Access 1 Repo

TL;DR

This paper introduces an adversarial testing methodology for Quality Estimation in Machine Translation, revealing that current models struggle with meaning errors and proposing a new way to evaluate their robustness without manual annotations.

Contribution

It presents a novel adversarial evaluation approach for QE systems, highlighting their limitations and proposing a predictive metric for model performance based on error detection ability.

Findings

01

State-of-the-art QE models still miss certain meaning errors.

02

Discrimination ability between meaning-preserving and altering perturbations correlates with overall QE performance.

03

Proposes a method to compare QE systems without manual quality annotations.

Abstract

Current Machine Translation (MT) systems achieve very good results on a growing variety of language pairs and datasets. However, they are known to produce fluent translation outputs that can contain important meaning errors, thus undermining their reliability in practice. Quality Estimation (QE) is the task of automatically assessing the performance of MT systems at test time. Thus, in order to be useful, QE systems should be able to detect such errors. However, this ability is yet to be tested in the current evaluation practices, where QE systems are assessed only in terms of their correlation with human judgements. In this work, we bridge this gap by proposing a general methodology for adversarial testing of QE for MT. First, we show that despite a high correlation with human judgements achieved by the recent SOTA, certain types of meaning errors are still problematic for QE to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

dipteshkanojia/qe-evaluation
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Authorship Attribution and Profiling

MethodsTest