How to Steal Reasoning Without Reasoning Traces

Tingwei Zhang; John X. Morris; Vitaly Shmatikov

arXiv:2603.07267·cs.CR·May 14, 2026

How to Steal Reasoning Without Reasoning Traces

Tingwei Zhang, John X. Morris, Vitaly Shmatikov

PDF

TL;DR

This paper introduces trace inversion models that generate detailed reasoning traces from limited model outputs, revealing reasoning capabilities and enabling model distillation.

Contribution

It presents a method to reconstruct detailed reasoning traces from black-box LLM outputs, facilitating model understanding and knowledge transfer.

Findings

01

Inverted traces closely match ground-truth reasoning when available.

02

Fine-tuning on inverted traces improves reasoning in student models.

03

Enables distillation from proprietary black-box LLMs.

Abstract

Many large language models (LLMs) use reasoning to generate responses but do not reveal their full reasoning traces (a.k.a. chains of thought), instead outputting only final answers and brief reasoning summaries. To demonstrate that hiding reasoning traces does not prevent users from "stealing" a model's reasoning capabilities, we introduce trace inversion models that, given only the inputs, answers, and (optionally) reasoning summaries exposed by a target model, generate detailed, synthetic reasoning traces. We show that (1) traces synthesized by trace inversion have high overlap with the ground-truth reasoning traces (when available), and (2) fine-tuning student models on inverted traces substantially improves their reasoning and enables distillation from proprietary, black-box LLMs.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.