Mangosteen: An Open Thai Corpus for Language Model Pretraining

Wannaphong Phatthiyaphaibun; Can Udomcharoenchaikit; Pakpoom Singkorapoom; Kunat Pipatanakul; Ekapol Chuangsuwanich; Peerat Limkonchotiwat; Sarana Nutanong

arXiv:2507.14664·cs.CL·July 23, 2025

Mangosteen: An Open Thai Corpus for Language Model Pretraining

Wannaphong Phatthiyaphaibun, Can Udomcharoenchaikit, Pakpoom Singkorapoom, Kunat Pipatanakul, Ekapol Chuangsuwanich, Peerat Limkonchotiwat, Sarana Nutanong

PDF

Open Access 2 Models 2 Datasets

TL;DR

Mangosteen is a comprehensive, openly available 47-billion-token Thai corpus created with a specialized pipeline, significantly improving Thai language model pretraining and setting a reproducible standard for future regional NLP research.

Contribution

The paper introduces Mangosteen, a large, transparent Thai corpus built with a Thai-adapted pipeline, and demonstrates its effectiveness in enhancing Thai language models.

Findings

01

Pipeline reduces web data from 202M to 25M documents.

02

Pretraining on Mangosteen improves Thai NLP benchmark scores.

03

Open release of data and tools supports reproducibility.

Abstract

Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnostic pipelines whose heuristics do not capture Thai script or cultural nuances, leaving risky material such as gambling content untreated. Prior Thai-specific efforts customize pipelines or build new ones, yet seldom release their data or document design choices, hindering reproducibility and raising the question of how to construct a transparent, high-quality Thai corpus. We introduce Mangosteen: a 47 billion-token Thai corpus built through a Thai-adapted Dolma pipeline that includes custom rule-based language ID, revised C4/Gopher quality filters, and Thai-trained content filters, plus curated non-web sources such as Wikipedia, Royal Gazette texts, OCR-extracted books, and CC-licensed YouTube subtitles.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

Datasets

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Computational and Text Analysis Methods · Topic Modeling