TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification

Adam Rida

arXiv:2604.14531·cs.AI·April 17, 2026

TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification

Adam Rida

PDF

1 Repo

TL;DR

TRACER is an open-source system that uses production logs to train surrogates for LLM classification, optimizing when to deploy them based on agreement thresholds to reduce inference costs.

Contribution

It introduces a trace-based adaptive routing system that determines surrogate deployment boundaries and provides interpretability artifacts for LLM classification tasks.

Findings

01

Achieves 83-100% surrogate coverage on intent benchmarks.

02

Fully replaces the teacher on a 150-class benchmark.

03

Correctly refuses deployment when representations are unreliable.

Abstract

Every call to an LLM classification endpoint produces a labeled input-output pair already retained in production logs. These pairs constitute a free, growing training set: a lightweight surrogate trained on them can absorb a significant portion of future traffic at near-zero marginal inference cost. The open questions are when the surrogate is reliable enough to deploy, what it handles versus defers, and how that boundary evolves as data accumulates. We introduce TRACER (Trace-based Adaptive Cost-Efficient Routing), an open-source system that trains ML surrogates on an LLM's own production traces and governs deployment through a parity gate: the surrogate is activated only when its agreement with the LLM exceeds a user-specified threshold {\alpha}. To make the routing boundary transparent, TRACER generates interpretability artifacts describing which input regions the surrogate…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

adrida/tracer
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.