Benchmarking Large Language Models for Persian: A Preliminary Study   Focusing on ChatGPT

Amirhossein Abaskohi; Sara Baruni; Mostafa Masoudi; Nesa Abbasi,; Mohammad Hadi Babalou; Ali Edalat; Sepehr Kamahi; Samin Mahdizadeh Sani,; Nikoo Naghavian; Danial Namazifard; Pouya Sadeghi; Yadollah Yaghoobzadeh

arXiv:2404.02403·cs.CL·April 4, 2024·2 cites

Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT

Amirhossein Abaskohi, Sara Baruni, Mostafa Masoudi, Nesa Abbasi,, Mohammad Hadi Babalou, Ali Edalat, Sepehr Kamahi, Samin Mahdizadeh Sani,, Nikoo Naghavian, Danial Namazifard, Pouya Sadeghi, Yadollah Yaghoobzadeh

PDF

Open Access 1 Repo

TL;DR

This study benchmarks large language models like GPT-3.5, GPT-4, and OpenChat-3.5 on Persian language tasks, revealing their strengths in reasoning and knowledge but also highlighting areas needing improvement, especially compared to fine-tuned models.

Contribution

First comprehensive benchmarking of LLMs on Persian, including new reasoning benchmarks and insights into performance differences with fine-tuned models.

Findings

01

GPT-4 excels in reasoning and knowledge tasks.

02

Translation to English improves GPT-3.5 performance.

03

Fine-tuned models outperform LLMs on specific tasks.

Abstract

This paper explores the efficacy of large language models (LLMs) for Persian. While ChatGPT and consequent LLMs have shown remarkable performance in English, their efficiency for more low-resource languages remains an open question. We present the first comprehensive benchmarking study of LLMs across diverse Persian language tasks. Our primary focus is on GPT-3.5-turbo, but we also include GPT-4 and OpenChat-3.5 to provide a more holistic evaluation. Our assessment encompasses a diverse set of tasks categorized into classic, reasoning, and knowledge-based domains. To enable a thorough comparison, we evaluate LLMs against existing task-specific fine-tuned models. Given the limited availability of Persian datasets for reasoning tasks, we introduce two new benchmarks: one based on elementary school math questions and another derived from the entrance exams for 7th and 10th grades. Our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ipouyall/benchmarking_chatgpt_for_persian
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsArtificial Intelligence in Healthcare and Education · Topic Modeling

MethodsRefunds@Expedia|||How do I get a full refund from Expedia? · {Dispute@FaQ-s}How to file a dispute with Expedia? · Attention Is All You Need · Sparse Evolutionary Training · Linear Layer · Layer Normalization · Multi-Head Attention · Weight Decay · Adam · Cosine Annealing