Advancing Bangla Machine Translation Through Informal Datasets
Ayon Roy, Risat Rahaman, Sadat Shibly, Udoy Saha Joy, Abdulla Al Kafi, Farig Yousuf Sadeque

TL;DR
This paper addresses the lack of informal Bangla-English translation datasets by developing new data from social media and conversations, aiming to improve open-source Bangla machine translation for everyday language.
Contribution
The paper introduces a novel dataset of informal Bangla-English pairs and proposes model improvements tailored for informal language translation.
Findings
Enhanced translation accuracy for informal Bangla.
Better accessibility of online information for Bangla speakers.
Foundation for future research in informal language translation.
Abstract
Bangla is the sixth most widely spoken language globally, with approximately 234 million native speakers. However, progress in open-source Bangla machine translation remains limited. Most online resources are in English and often remain untranslated into Bangla, excluding millions from accessing essential information. Existing research in Bangla translation primarily focuses on formal language, neglecting the more commonly used informal language. This is largely due to the lack of pairwise Bangla-English data and advanced translation models. If datasets and models can be enhanced to better handle natural, informal Bangla, millions of people will benefit from improved online information access. In this research, we explore current state-of-the-art models and propose improvements to Bangla translation by developing a dataset from informal sources like social media and conversational…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsNatural Language Processing Techniques · ICT in Developing Communities · Translation Studies and Practices
