Awal -- Community-Powered Language Technology for Tamazight
Alp \"Oktem, Farida Boudichat

TL;DR
Awal is a community-driven platform launched in 2024 to develop Tamazight language technology resources, addressing data scarcity through collaborative efforts despite sociolinguistic challenges.
Contribution
This paper introduces Awal, a novel community-powered initiative that leverages local speaker contributions to build NLP resources for Tamazight, highlighting challenges and initial results.
Findings
Community engagement faced barriers like low confidence and standardization issues.
Data contributions were limited, with 6,421 translation pairs and 3 hours of speech.
Positive reception but limited participation from non-linguists.
Abstract
This paper presents Awal, a community-powered initiative for developing language technology resources for Tamazight. We provide a comprehensive review of the NLP landscape for Tamazight, examining recent progress in computational resources, and the emergence of community-driven approaches to address persistent data scarcity. Launched in 2024, awaldigital.org platform addresses the underrepresentation of Tamazight in digital spaces through a collaborative platform enabling speakers to contribute translation and voice data. We analyze 18 months of community engagement, revealing significant barriers to participation including limited confidence in written Tamazight and ongoing standardization challenges. Despite widespread positive reception, actual data contribution remained concentrated among linguists and activists. The modest scale of community contributions -- 6,421 translation pairs…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsICT in Developing Communities · Mobile Crowdsensing and Crowdsourcing · Natural Language Processing Techniques
