BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing   Critical Translation Errors in Sentiment-oriented Text

Hadeel Saadany; Constantin Orasan

arXiv:2109.14250·cs.CL·September 30, 2021

BLEU, METEOR, BERTScore: Evaluation of Metrics Performance in Assessing Critical Translation Errors in Sentiment-oriented Text

Hadeel Saadany, Constantin Orasan

PDF

TL;DR

This paper evaluates how well common automatic translation quality metrics detect critical errors in sentiment-oriented social media content, highlighting the need for improved metrics to ensure accurate sentiment translation.

Contribution

It compares the effectiveness of three standard metrics in identifying sentiment-critical translation errors, revealing their limitations and the need for fine-tuning.

Findings

01

Metrics struggle to detect sentiment-critical errors in translations.

02

Current metrics are insufficient for reliable sentiment error detection.

03

Fine-tuning is necessary to improve robustness of evaluation metrics.

Abstract

Social media companies as well as authorities make extensive use of artificial intelligence (AI) tools to monitor postings of hate speech, celebrations of violence or profanity. Since AI software requires massive volumes of data to train computers, Machine Translation (MT) of the online content is commonly used to process posts written in several languages and hence augment the data needed for training. However, MT mistakes are a regular occurrence when translating sentiment-oriented user-generated content (UGC), especially when a low-resource language is involved. The adequacy of the whole process relies on the assumption that the evaluation metrics used give a reliable indication of the quality of the translation. In this paper, we assess the ability of automatic quality metrics to detect critical machine translation errors which can cause serious misunderstanding of the affect…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.