ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman and, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid, Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav, Nakov, Timothy Baldwin

TL;DR
This paper introduces ArabicMMLU, a comprehensive multi-task benchmark for evaluating Arabic language understanding across diverse educational tasks, revealing significant performance gaps in current models.
Contribution
It presents the first multi-task benchmark for Arabic, sourced from real school exams, enabling systematic evaluation of language models in Arabic.
Findings
Most models score below 50% on the benchmark.
Top models only reach around 62.3% accuracy.
Significant room for improvement in Arabic NLP models.
Abstract
The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present \datasetname{}, the first multi-task language understanding benchmark for the Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA) and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
- 🤗inceptionai/jais-adapted-70b-chatmodel· 54 dl· ♡ 954 dl♡ 9
- 🤗inceptionai/jais-adapted-70bmodel· 144 dl· ♡ 21144 dl♡ 21
- 🤗inceptionai/jais-family-590m-chatmodel· 34 dl· ♡ 734 dl♡ 7
- 🤗inceptionai/jais-family-590mmodel· 404 dl· ♡ 7404 dl♡ 7
- 🤗inceptionai/jais-adapted-7bmodel· 837 dl· ♡ 8837 dl♡ 8
- 🤗inceptionai/jais-family-1p3b-chatmodel· 164 dl· ♡ 6164 dl♡ 6
- 🤗inceptionai/jais-adapted-7b-chatmodel· 1.7k dl· ♡ 81.7k dl♡ 8
- 🤗inceptionai/jais-family-1p3bmodel· 12 dl· ♡ 1012 dl♡ 10
- 🤗inceptionai/jais-adapted-13bmodel· ♡ 5♡ 5
- 🤗inceptionai/jais-adapted-13b-chatmodel· 2.4k dl· ♡ 72.4k dl♡ 7
Videos
Taxonomy
TopicsNatural Language Processing Techniques · Text Readability and Simplification · Topic Modeling
MethodsFocus · mT0 · BLOOMZ
