Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap

Zijun Wang; Yijiahao Qi; Hanqiu Chen; Zishen Wan; Gongjin Sun; Dongyang Li; Shuyi Pei; Cong Hao

arXiv:2512.18126·cs.AI·December 23, 2025

Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap

Zijun Wang, Yijiahao Qi, Hanqiu Chen, Zishen Wan, Gongjin Sun, Dongyang Li, Shuyi Pei, Cong Hao

PDF

Open Access

TL;DR

This paper introduces a novel MoA serving approach that reduces latency by restructuring agent communication, adaptively skipping computations, and overlapping processing, achieving up to 90% latency reduction with maintained accuracy.

Contribution

It proposes a hierarchical tree topology, adaptive runtime skipping, and overlapping execution techniques to improve MoA serving efficiency and latency.

Findings

01

Up to 90% reduction in end-to-end latency.

02

Maintains accuracy within ±1% of dense MoA baselines.

03

Improves hardware utilization and inference speed.

Abstract

Mixture-of-Agents (MoA) inference can suffer from dense inter-agent communication and low hardware utilization, which jointly inflate serving latency. We present a serving design that targets these bottlenecks through an algorithm-system co-design. First, we replace dense agent interaction graphs with a hierarchical tree topology that induces structured sparsity in inter-agent communication. Second, we introduce a runtime adaptive mechanism that selectively terminates or skips downstream agent invocations using semantic agreement and confidence signals from intermediate outputs. Third, we pipeline agent execution by overlapping incremental prefilling with decoding across dependency-related agents, improving utilization and reducing inference latency. Across representative tasks, this approach substantially reduces end-to-end latency (up to 90%) while maintaining comparable accuracy…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Machine Learning and Algorithms · Data Stream Mining Techniques