ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models

Yingjie Xia; Tao Liu; Jinglei Shi; Qingsong Xie; Heng Guo; Jian Yang; Xi Wang

arXiv:2512.05385·cs.CV·March 17, 2026

ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models

Yingjie Xia, Tao Liu, Jinglei Shi, Qingsong Xie, Heng Guo, Jian Yang, Xi Wang

PDF

Open Access

TL;DR

ShaRP introduces a novel shallow-layer pruning method for Video Large Language Models, addressing unreliable attention scores in early layers to significantly reduce computational costs while maintaining high performance.

Contribution

The paper identifies a failure mode in shallow-layer attention pruning and proposes ShaRP, a unified framework that improves token selection reliability for efficient VLLM inference.

Findings

01

Preserves 97.2% of original performance

02

Reduces TFLOPs by 86%

03

Achieves 5.1x speedup in prefilling stage

Abstract

Video Large Language Models (VLLMs) incur substantial prefilling cost due to the large number of visual tokens. While attention-based token pruning offers a promising acceleration strategy, applying it at shallow decoder layers often causes severe performance degradation under high compression ratios, limiting its practical benefits. In this work, we uncover an overlooked failure mode in shallow-layer attention pruning: attention scores in early decoder layers can become unreliable indicators of token utility, resulting in unstable token selection under aggressive compression. We show that this effect arises from the joint influence of insufficient token interaction, content-agnostic positional bias, and redundancy among high-attention tokens, which together distort attention-based importance estimation before informative representations fully emerge. Motivated by this insight, we…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Generative Adversarial Networks and Image Synthesis · Domain Adaptation and Few-Shot Learning