Runtime Burden Allocation for Structured LLM Routing in Agentic Expert Systems: A Full-Factorial Cross-Backend Methodology

Zhou Hanlin; Chan Huah Yong

arXiv:2604.01235·cs.AI·April 3, 2026

Runtime Burden Allocation for Structured LLM Routing in Agentic Expert Systems: A Full-Factorial Cross-Backend Methodology

Zhou Hanlin, Chan Huah Yong

PDF

TL;DR

This paper introduces a systems-level approach to structured LLM routing, emphasizing the importance of backend-specific strategies to optimize correctness, latency, and cost in agentic AI systems.

Contribution

It presents a comprehensive full-factorial benchmark and a deployable framework for evaluating and optimizing structured LLM routing across diverse backends.

Findings

01

No universal best routing mode; backend-specific effects dominate performance.

02

Reliability varies significantly across backends like Gemini, OpenAI, and Llama.

03

Efficiency gains from compression are highly backend-dependent.

Abstract

Structured LLM routing is often treated as a prompt-engineering problem. We argue that it is, more fundamentally, a systems-level burden-allocation problem. As large language models (LLMs) become core control components in agentic AI systems, reliable structured routing must balance correctness, latency, and implementation cost under real deployment constraints. We show that this balance is shaped not only by prompts or schemas, but also by how structural work is allocated across the generation stack: whether output structure is emitted directly by the model, compressed during transport, or reconstructed locally after generation. We evaluate this formulation through a comprehensive full-factorial benchmark covering 48 deployment configurations and 15,552 requests across OpenAI, Gemini, and Llama backends. Our central finding is consequential: there is no universal best routing mode.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.