ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
Neha Joshi, Pamir Gogoi, Aasim Mirza, Aayush Jansari, Aditya Yadavalli, Ayushi Pandey, Arunima Shukla, Deepthi Sudharsan, Kalika Bali, Vivek Seshadri

TL;DR
This paper introduces ELR-1000, a multimodal dataset of 1,060 traditional recipes in 10 endangered Indian languages, and evaluates language models' translation performance, highlighting challenges and improvements with contextual information.
Contribution
The paper presents a new culturally-grounded dataset for endangered languages and analyzes how contextual prompts improve translation quality in low-resource settings.
Findings
Large language models struggle with low-resource, culturally-specific languages.
Providing targeted context significantly improves translation accuracy.
The dataset encourages development of equitable language technologies.
Abstract
We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural nuance, were collected using a mobile interface designed for contributors with low digital literacy. Endangered Language Recipes (ELR)-1000 -- captures not only culinary practices but also the socio-cultural context embedded in indigenous food traditions. We evaluate the performance of several state-of-the-art large language models (LLMs) on translating these recipes into English and find the following: despite the models' capabilities, they struggle with low-resource, culturally-specific language. However, we observe that providing targeted context -- including background information about the languages, translation examples, and guidelines for…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsICT in Developing Communities · Language and cultural evolution · Multilingual Education and Policy
