Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

Cai Zhou; Zekai Wang; Menghua Wu; Qianyu Julie Zhu; Flora C. Shi; Chenyu Wang; Ashia Wilson; Tommi Jaakkola; Stephen Bates

arXiv:2604.01170·cs.LG·April 2, 2026

Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi, Chenyu Wang, Ashia Wilson, Tommi Jaakkola, Stephen Bates

PDF

1 Repo 1 Models 1 Datasets

TL;DR

ORCA introduces a test-time calibration framework for large language models that improves reasoning efficiency and generalization by dynamically calibrating sampling processes with conformal prediction and meta-learning.

Contribution

It presents a novel online calibration method that provides theoretical guarantees and enhances reasoning task performance across diverse settings.

Findings

01

Up to 47.5% efficiency savings on in-distribution tasks.

02

67.0% savings in out-of-domain zero-shot settings.

03

Maintains low empirical error rates across benchmarks.

Abstract

While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language models, and the lack of calibration in popular sampling techniques. Here, we present Online Reasoning Calibration (ORCA), a framework for calibrating the sampling process that draws upon conformal prediction and test-time training. Specifically, we introduce a meta-learning procedure that updates the calibration module for each input. This allows us to provide valid confidence estimates under distributional shift, e.g. in thought patterns that occur across different stages of reasoning, or in prompt distributions between model development and deployment. ORCA not only provides theoretical guarantees on conformal risks, but also empirically shows higher…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

wzekai99/ORCA
github

Models

🤗
wzekai99/ORCA
model

Datasets

wzekai99/ORCA
dataset· 138 dl
138 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.