LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference

Shashank Kapadia; Deep Naryan Mishra; Sujal Reddy Alugubelli; Haoan Wang; Saipraveen Vabbilisetty; Rishi Bhatia; Anupriya Sharma

arXiv:2605.01058·cs.LG·May 5, 2026

LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference

Shashank Kapadia, Deep Naryan Mishra, Sujal Reddy Alugubelli, Haoan Wang, Saipraveen Vabbilisetty, Rishi Bhatia, Anupriya Sharma

PDF

TL;DR

LEAP introduces a new pretraining method that aligns intermediate transformer layers with final representations, enabling efficient early exit inference without architectural changes.

Contribution

It proposes LEAP, a training objective that reconciles distillation and early exit, improving inference speedup while maintaining performance.

Findings

01

LEAP achieves 1.61× wall-clock speedup on NVIDIA L4 hardware.

02

91.9% of samples exit by layer 7 with LEAP-MiniLM.

03

LEAP provides effective layer reduction and operational deployment guidance.

Abstract

Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; yet we establish that they exhibit systematic incompatibility under standard deployment conditions for convergence-based early exit. Distillation objectives that align intermediate student layers to teacher representations suppress the representational convergence that early-exit mechanisms exploit, rendering such mechanisms ineffective on distilled models. We introduce LEAP (Layer-wise Exit-Aware Pretraining), an auxiliary training objective that reconciles this incompatibility. LEAP requires no architectural modifications; it augments standard distillation with a single constraint ensuring intermediate layers approximate final-layer representations. LEAP-MiniLM achieves 1.61 $\times$ measured wall-clock speedup (batch=1, NVIDIA L4) at…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.