IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, Partha, Talukdar

TL;DR
IndicGenBench is a comprehensive multilingual benchmark designed to evaluate large language models on generation tasks across 29 Indic languages, highlighting performance gaps and the need for more inclusive models.
Contribution
The paper introduces IndicGenBench, the largest benchmark for Indic languages, with diverse tasks and human-curated data, enabling evaluation of LLMs on under-represented languages for the first time.
Findings
PaLM-2 performs best among evaluated models.
Significant performance gap exists between Indic languages and English.
Further research needed for inclusive multilingual models.
Abstract
As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world. India is a linguistically diverse country of 1.4 Billion people. To facilitate research on multilingual LLM evaluation, we release IndicGenBench - the largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts and 4 language families. IndicGenBench is composed of diverse generation tasks like cross-lingual summarization, machine translation, and cross-lingual question answering. IndicGenBench extends existing benchmarks to many Indic languages through human curation providing multi-way parallel evaluation data for many under-represented Indic languages for the first time. We evaluate a wide range of proprietary and open-source LLMs including GPT-3.5,…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
- 🤗google/gemma-3-4b-itmodel· 1.5M dl· ♡ 12721.5M dl♡ 1272
- 🤗google/gemma-3-27b-itmodel· 1.0M dl· ♡ 19401.0M dl♡ 1940
- 🤗unsloth/gemma-3-12b-it-GGUFmodel· 101k dl· ♡ 178101k dl♡ 178
- 🤗google/gemma-3-1b-itmodel· 1.4M dl· ♡ 8991.4M dl♡ 899
- 🤗google/gemma-3-12b-it-qat-q4_0-ggufmodel· 7.1k dl· ♡ 2627.1k dl♡ 262
- 🤗google/gemma-3-270mmodel· 83k dl· ♡ 100383k dl♡ 1003
- 🤗google/gemma-3-12b-itmodel· 2.6M dl· ♡ 6982.6M dl♡ 698
- 🤗google/gemma-3-12b-it-qat-q4_0-unquantizedmodel· 28k dl· ♡ 8128k dl♡ 81
- 🤗p-e-w/gemma-3-12b-it-hereticmodel· 2.4k dl· ♡ 792.4k dl♡ 79
- 🤗llmfan46/gemma-3-12b-it-ultra-uncensored-heretic-GGUFmodel· 23k dl· ♡ 1323k dl♡ 13
Videos
Taxonomy
TopicsNatural Language Processing Techniques · Translation Studies and Practices
MethodsRefunds@Expedia|||How do I get a full refund from Expedia? · 15 Ways to Contact How can i speak to someone at Delta Airlines · Attention Is All You Need · Sparse Evolutionary Training · Adafactor · Position-Wise Feed-Forward Layer · SentencePiece · Inverse Square Root Schedule · Absolute Position Encodings · Linear Layer
