Efficient Contextual LLM Cascades through Budget-Constrained Policy   Learning

Xuechen Zhang; Zijian Huang; Ege Onur Taga; Carlee Joe-Wong; Samet; Oymak; Jiasi Chen

arXiv:2404.13082·cs.CL·November 21, 2024

Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning

Xuechen Zhang, Zijian Huang, Ege Onur Taga, Carlee Joe-Wong, Samet, Oymak, Jiasi Chen

PDF

Open Access 1 Video

TL;DR

This paper introduces TREACLE, a reinforcement learning policy that optimally selects among multiple large language models and prompts to minimize costs while meeting accuracy, latency, and budget constraints in question-answering tasks.

Contribution

The paper presents TREACLE, a novel RL-based approach for dynamic model and prompt selection that considers context and constraints, improving cost efficiency in LLM inference.

Findings

01

Achieves up to 85% cost savings compared to baselines.

02

Maintains high accuracy while reducing costs.

03

Enables flexible trade-offs between accuracy and expense.

Abstract

Recent successes in natural language processing have led to the proliferation of large language models (LLMs) by multiple providers. Each LLM offering has different inference accuracy, monetary cost, and latency, and their accuracy further depends on the exact wording of the question (i.e., the specific prompt). At the same time, users often have a limit on monetary budget and latency to answer all their questions, and they do not know which LLMs to choose for each question to meet their accuracy and long term budget requirements. To navigate this rich design space, we propose TREACLE ( $\underline{T}$ hrifty $\underline{R e a}$ soning via $\underline{C}$ ontext-Aware $\underline{L}$ LM and Prompt S $\underline{e}$ lection), a reinforcement learning policy that jointly selects the model and prompting scheme while respecting the user's monetary cost and latency constraints. TREACLE uses the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning· slideslive

Taxonomy

TopicsAnomaly Detection Techniques and Applications · Time Series Analysis and Forecasting · Machine Learning and Data Classification