Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

Hunzalah Hassan Bhatti; Firoj Alam

arXiv:2510.24328·cs.CL·April 20, 2026

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

Hunzalah Hassan Bhatti, Firoj Alam

PDF

1 Datasets

TL;DR

This paper introduces a new Arabic cultural QA benchmark with dialect variants, translating questions into multiple dialects, converting them into open-ended formats, and evaluating LLMs' performance with chain-of-thought reasoning.

Contribution

It presents the first parallel dataset of Arabic QA across dialects, extending existing datasets, and evaluates LLMs' performance on culturally grounded, dialect-specific questions.

Findings

01

Models underperform on Arabic dialects, especially on open-ended questions.

02

Arabic-centric models excel at MCQs but struggle with OEQs.

03

Chain-of-thought reasoning improves correctness judgments but affects n-gram metrics.

Abstract

Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across multiple language varieties, making it, to our knowledge, the first of its kind. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Datasets

QCRI/ArabicCulturalQA
dataset· 209 dl
209 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.