Exploiting Out-of-Domain Data Sources for Dialectal Arabic Statistical   Machine Translation

Katrin Kirchhoff; Bing Zhao; Wen Wang

arXiv:1509.01938·cs.CL·September 8, 2015

Exploiting Out-of-Domain Data Sources for Dialectal Arabic Statistical Machine Translation

Katrin Kirchhoff, Bing Zhao, Wen Wang

PDF

Open Access

TL;DR

This paper presents methods to extract dialect-specific parallel data from out-of-domain Arabic corpora to improve statistical machine translation for Iraqi Arabic, demonstrating that targeted data selection enhances translation quality.

Contribution

It introduces data selection techniques for dialectal Arabic MT and explores using automatically translated speech data as additional training material.

Findings

01

Targeted data selection improves translation performance

02

Small, highly relevant datasets are effective

03

Preliminary results show promise for speech data integration

Abstract

Statistical machine translation for dialectal Arabic is characterized by a lack of data since data acquisition involves the transcription and translation of spoken language. In this study we develop techniques for extracting parallel data for one particular dialect of Arabic (Iraqi Arabic) from out-of-domain corpora in different dialects of Arabic or in Modern Standard Arabic. We compare two different data selection strategies (cross-entropy based and submodular selection) and demonstrate that a very small but highly targeted amount of found data can improve the performance of a baseline machine translation system. We furthermore report on preliminary experiments on using automatically translated speech data as additional training data.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Machine Learning and Algorithms