Ministral 3

Alexander H. Liu; Kartik Khandelwal; Sandeep Subramanian; Victor Jouault; Abhinav Rastogi; Adrien Sad\'e; Alan Jeffares; Albert Jiang; Alexandre Cahill; Alexandre Gavaudan; Alexandre Sablayrolles; Am\'elie H\'eliou; Amos You; Andy Ehrenberg; Andy Lo; Anton Eliseev; Antonia Calvi; Avinash Sooriyarachchi; Baptiste Bout; Baptiste Rozi\`ere; Baudouin De Monicault; Cl\'emence Lanfranchi; Corentin Barreau; Cyprien Courtot; Daniele Grattarola; Darius Dabert; Diego de las Casas; Elliot Chane-Sane; Faruk Ahmed; Gabrielle Berrada; Ga\"etan Ecrepont; Gauthier Guinet; Georgii Novikov; Guillaume Kunsch; Guillaume Lample; Guillaume Martin; Gunshi Gupta; Jan Ludziejewski; Jason Rute; Joachim Studnia; Jonas Amar; Jos\'ephine Delas; Josselin Somerville Roberts; Karmesh Yadav; Khyathi Chandu; Kush Jain; Laurence Aitchison; Laurent Fainsin; L\'eonard Blier; Lingxiao Zhao; Louis Martin; Lucile Saulnier; Luyu Gao; Maarten Buyl; Margaret Jennings; Marie Pellat; Mark Prins; Mathieu Poir\'ee; Mathilde Guillaumin; Matthieu Dinot; Matthieu Futeral; Maxime Darrin; Maximilian Augustin; Mia Chiquier; Michel Schimpf; Nathan Grinsztajn; Neha Gupta; Nikhil Raghuraman; Olivier Bousquet; Olivier Duchenne; Patricia Wang; Patrick von Platen; Paul Jacob; Paul Wambergue; Paula Kurylowicz; Pavankumar Reddy Muddireddy; Philom\`ene Chagniot; Pierre Stock; Pravesh Agrawal; Quentin Torroba; Romain Sauvestre; Roman Soletskyi; Rupert Menneer; Sagar Vaze; Samuel Barry; Sanchit Gandhi; Siddhant Waghjale; Siddharth Gandhi; Soham Ghosh; Srijan Mishra; Sumukh Aithal; Szymon Antoniak; Teven Le Scao; Th\'eo Cachet; Theo Simon Sorg; Thibaut Lavril; Thiziri Nait Saada; Thomas Chabal; Thomas Foubert; Thomas Robert; Thomas Wang; Tim Lawson; Tom Bewley; Tom Bewley; Tom Edwards; Umar Jamil; Umberto Tomasini; Valeriia Nemychnikova; Van Phung; Vincent Maladi\`ere; Virgile Richard; Wassim Bouaziz; Wen-Ding Li; William Marshall; Xinghui Li; Xinyu Yang; Yassine El Ouahidi; Yihan Wang; Yunhao Tang; Zaccharie Ramzi

arXiv:2601.08584·cs.CL·January 14, 2026

Ministral 3

Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sad\'e, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, Alexandre Sablayrolles, Am\'elie H\'eliou, Amos You, Andy Ehrenberg, Andy Lo, Anton Eliseev, Antonia Calvi

PDF

Open Access 10 Models 1 Datasets

TL;DR

Minstral 3 introduces a family of parameter-efficient dense language models in three sizes, optimized for constrained environments, with variants for general use, instruction tuning, and reasoning, incorporating image understanding and Cascade Distillation.

Contribution

The paper presents Ministral 3, a new series of dense language models with multiple variants and a novel Cascade Distillation method for efficient model derivation.

Findings

01

Models are available in 3B, 8B, and 14B sizes.

02

All models include image understanding capabilities.

03

Models are released under Apache 2.0 license.

Abstract

We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes: 3B, 8B, and 14B parameters. For each model size, we release three variants: a pretrained base model for general-purpose use, an instruction finetuned, and a reasoning model for complex problem-solving. In addition, we present our recipe to derive the Ministral 3 models through Cascade Distillation, an iterative pruning and continued training with distillation technique. Each model comes with image understanding capabilities, all under the Apache 2.0 license.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

Datasets

dmenacho/Fatima_DMO_application
dataset· 55 dl
55 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Natural Language Processing Techniques · Topic Modeling