Statistical Machine Translation for Indian Languages: Mission Hindi 2

Raj Nath Patel; Prakash B. Pimpale

arXiv:1610.08000·cs.CL·October 26, 2016

Statistical Machine Translation for Indian Languages: Mission Hindi 2

Raj Nath Patel, Prakash B. Pimpale

PDF

Open Access

TL;DR

This paper details CDAC Mumbai's approach to statistical machine translation for five Indian language pairs across multiple domains, employing preprocessing techniques like suffix separation, compound splitting, and preordering to improve translation quality.

Contribution

The paper introduces a comprehensive SMT system for five Indian language pairs, applying specific preprocessing techniques to enhance translation effectiveness across diverse domains.

Findings

01

Effective translation for all language pairs tested.

02

Preprocessing techniques improved translation quality.

03

System demonstrated robustness across domains.

Abstract

This paper presents Centre for Development of Advanced Computing Mumbai's (CDACM) submission to NLP Tools Contest on Statistical Machine Translation in Indian Languages (ILSMT) 2015 (collocated with ICON 2015). The aim of the contest was to collectively explore the effectiveness of Statistical Machine Translation (SMT) while translating within Indian languages and between English and Indian languages. In this paper, we report our work on all five language pairs, namely Bengali-Hindi (\bnhi), Marathi-Hindi (\mrhi), Tamil-Hindi (\tahi), Telugu-Hindi (\tehi), and English-Hindi (\enhi) for Health, Tourism, and General domains. We have used suffix separation, compound splitting and preordering prior to SMT training and testing.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques