DeepSubQE: Quality estimation for subtitle translations

Prabhakar Gupta; Anil Nelakanti

arXiv:2004.13828·cs.CL·April 30, 2020·1 cites

DeepSubQE: Quality estimation for subtitle translations

Prabhakar Gupta, Anil Nelakanti

PDF

Open Access

TL;DR

DeepSubQE is a novel system for estimating the quality of video subtitle translations, leveraging hybrid neural networks and data augmentation to improve accuracy over existing methods.

Contribution

The paper introduces DeepSubQE, a hybrid neural network model with data augmentation strategies specifically designed for subtitle translation quality estimation.

Findings

01

DeepSubQE outperforms existing QE methods significantly.

02

Hybrid network combining semantic and syntactic features is more effective.

03

Data augmentation improves training and model robustness.

Abstract

Quality estimation (QE) for tasks involving language data is hard owing to numerous aspects of natural language like variations in paraphrasing, style, grammar, etc. There can be multiple answers with varying levels of acceptability depending on the application at hand. In this work, we look at estimating quality of translations for video subtitles. We show how existing QE methods are inadequate and propose our method DeepSubQE as a system to estimate quality of translation given subtitles data for a pair of languages. We rely on various data augmentation strategies for automated labelling and synthesis for training. We create a hybrid network which learns semantic and syntactic features of bilingual data and compare it with only-LSTM and only-CNN networks. Our proposed network outperforms them by significant margin.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Subtitles and Audiovisual Media · Multimodal Machine Learning Applications