Efficient Reasoning on the Edge

Yelysei Bondarenko; Thomas Hehn; Rob Hesselink; Romain Lepert; Fabio Valerio Massoli; Evgeny Mironov; Leyla Mirvakhabova; Tribhuvanesh Orekondy; Spyridon Stasis; Andrey Kuzmin; Anna Kuzina; Markus Nagel; Ankita Nayak; Corrado Rainone; Ork de Rooij; Paul N Whatmough; Arash Behboodi; Babak Ehteshami Bejnordi

arXiv:2603.16867·cs.LG·March 18, 2026

Efficient Reasoning on the Edge

Yelysei Bondarenko, Thomas Hehn, Rob Hesselink, Romain Lepert, Fabio Valerio Massoli, Evgeny Mironov, Leyla Mirvakhabova, Tribhuvanesh Orekondy, Spyridon Stasis, Andrey Kuzmin, Anna Kuzina, Markus Nagel, Ankita Nayak, Corrado Rainone, Ork de Rooij, Paul N Whatmough

PDF

Open Access

TL;DR

This paper introduces a lightweight, resource-efficient method for enabling reasoning in small language models suitable for mobile devices, using LoRA adapters, reinforcement learning, and dynamic mechanisms to reduce costs and maintain accuracy.

Contribution

The authors propose a novel approach combining LoRA adapters, reinforcement learning, and dynamic adapter switching to enable efficient reasoning in small LLMs for edge deployment.

Findings

01

Achieves significant reduction in response length with minimal accuracy loss.

02

Improves on-device reasoning efficiency using memory and computation optimizations.

03

Demonstrates practical reasoning capabilities on mobile devices with Qwen2.5-7B.

Abstract

Large language models (LLMs) with chain-of-thought reasoning achieve state-of-the-art performance across complex problem-solving tasks, but their verbose reasoning traces and large context requirements make them impractical for edge deployment. These challenges include high token generation costs, large KV-cache footprints, and inefficiencies when distilling reasoning capabilities into smaller models for mobile devices. Existing approaches often rely on distilling reasoning traces from larger models into smaller models, which are verbose and stylistically redundant, undesirable for on-device inference. In this work, we propose a lightweight approach to enable reasoning in small LLMs using LoRA adapters combined with supervised fine-tuning. We further introduce budget forcing via reinforcement learning on these adapters, significantly reducing response length with minimal accuracy loss.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBig Data and Digital Economy · Advanced Neural Network Applications · Green IT and Sustainability