RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models

Jiarui Zhang; Xiangyu Liu; Yong Hu; Chaoyue Niu; Fan Wu; Guihai Chen

arXiv:2505.23052·cs.CL·October 20, 2025

RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models

Jiarui Zhang, Xiangyu Liu, Yong Hu, Chaoyue Niu, Fan Wu, Guihai Chen

PDF

Open Access

TL;DR

RAGRouter introduces a dynamic routing method for multiple retrieval-augmented language models, improving task performance by considering document influence and knowledge shifts, outperforming existing static routing approaches.

Contribution

It formally defines the retrieval-augmented LLM routing problem and proposes RAGRouter, a novel contrastive learning-based routing framework that adapts to document influence.

Findings

01

RAGRouter outperforms individual LLMs and existing routing methods.

02

It achieves a strong performance-efficiency trade-off under low-latency constraints.

03

Extensive experiments validate its effectiveness across diverse tasks and models.

Abstract

Retrieval-Augmented Generation (RAG) significantly improves the performance of Large Language Models (LLMs) on knowledge-intensive tasks. However, varying response quality across LLMs under RAG necessitates intelligent routing mechanisms, which select the most suitable model for each query from multiple retrieval-augmented LLMs via a dedicated router model. We observe that external documents dynamically affect LLMs' ability to answer queries, while existing routing methods, which rely on static parametric knowledge representations, exhibit suboptimal performance in RAG scenarios. To address this, we formally define the new retrieval-augmented LLM routing problem, incorporating the influence of retrieved documents into the routing framework. We propose RAGRouter, a RAG-aware routing design, which leverages document embeddings and RAG capability embeddings with contrastive learning to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Graph Neural Networks · Topic Modeling · Information Retrieval and Search Behavior

MethodsRefunds@Expedia|||How do I get a full refund from Expedia? · Attention Is All You Need · Linear Layer · Byte Pair Encoding · Attention Dropout · Softmax · WordPiece · BART · Weight Decay · Multi-Head Attention