MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

Lu Xu; Jiaqian Yu; Xiongfeng Peng; Yiwei Chen; Weiming Li; Jaewook Yoo; Sunghyun Chunag; Dongwook Lee; Daehyun Ji; Chao Zhang

arXiv:2507.07818·cs.AI·August 14, 2025

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

Lu Xu, Jiaqian Yu, Xiongfeng Peng, Yiwei Chen, Weiming Li, Jaewook Yoo, Sunghyun Chunag, Dongwook Lee, Daehyun Ji, Chao Zhang

PDF

TL;DR

MoSE introduces a skill-by-skill mixture-of-experts approach for embodied AI, improving reasoning and learning efficiency in autonomous systems by mimicking human skill learning and reasoning processes.

Contribution

The paper proposes a novel skill-oriented MoE model with a hierarchical dataset and routing mechanism, enabling efficient skill-by-skill learning in embodied AI tasks.

Findings

01

Outperforms existing models on autonomous driving reasoning tasks

02

Achieves better results with less than 40% of parameters

03

Effectively integrates auxiliary tasks without extra computational cost

Abstract

To meet the growing demand for smarter, faster, and more efficient embodied AI solutions, we introduce a novel Mixture-of-Expert (MoE) method that significantly boosts reasoning and learning efficiency for embodied autonomous systems. General MoE models demand extensive training data and complex optimization, which limits their applicability in embodied AI such as autonomous driving (AD) and robotic manipulation. In this work, we propose a skill-oriented MoE called MoSE, which mimics the human learning and reasoning process skill-by-skill, step-by-step. We introduce a skill-oriented routing mechanism that begins with defining and annotating specific skills, enabling experts to identify the necessary competencies for various scenarios and reasoning tasks, thereby facilitating skill-by-skill learning. To better align with multi-step planning in human reasoning and in end-to-end driving…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.