Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

Fan Jiang; Yu Zhao; Chenyang Lyu; Tianqi Shi; Yichao Du; Feihu Jiang; Longyue Wang; Weihua Luo

arXiv:2604.25578·cs.CL·April 29, 2026

Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

Fan Jiang, Yu Zhao, Chenyang Lyu, Tianqi Shi, Yichao Du, Feihu Jiang, Longyue Wang, Weihua Luo

PDF

TL;DR

Marco-MoE introduces a highly sparse, open multilingual mixture-of-experts model that achieves efficient training and superior performance, with scalable language expansion and shared activation patterns across languages.

Contribution

The paper presents Marco-MoE, a novel sparse multilingual MoE model that outperforms similar-sized models and enables scalable language expansion with shared expert activation patterns.

Findings

01

Achieves state-of-the-art performance-to-compute ratio.

02

Surpasses larger models in instruction tuning.

03

Learns shared activation patterns across related languages.

Abstract

We present Marco-MoE, a suite of fully open multilingual sparse Mixture-of-Experts (MoE) models. Marco-MoE features a highly sparse design in which only around 5\% of the total parameters are activated per input token. This extreme sparsity, combined with upcycling from dense models, enables efficient pre-training on 5T tokens. Our models surpass similarly-sized competitors on English and multilingual benchmarks, achieving a best-in-class performance-to-compute ratio. We further post-train these models to create Marco-MoE-\textsc{Instruct} variants, which surpass the performance of competing models possessing $3$ -- $14 \times$ more activated parameters. Our analysis reveals that Marco-MoE learns structured expert activation patterns shared across related languages, while maintaining highly specialized utilization for linguistically isolated ones. We further show that Marco-MoE allows for…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.