Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models

Chiyue Wei; Cong Guo; Junyao Zhang; Haoxuan Shan; Yifan Xu; Ziyue Zhang; Yudong Liu; Qinsi Wang; Changchun Zhou; Hai "Helen" Li; Yiran Chen

arXiv:2512.14661·cs.AR·December 17, 2025

Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models

Chiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan, Yifan Xu, Ziyue Zhang, Yudong Liu, Qinsi Wang, Changchun Zhou, Hai "Helen" Li, Yiran Chen

PDF

Open Access

TL;DR

Focus introduces a multilevel, streaming architecture for vision-language models that significantly reduces computation and energy consumption, enabling real-time inference on hardware accelerators.

Contribution

The paper presents a novel streaming concentration architecture with hierarchical redundancy elimination tailored for efficient vision-language model inference.

Findings

01

2.4x speedup in inference throughput

02

3.3x reduction in energy consumption

03

Outperforms state-of-the-art accelerators in efficiency

Abstract

Vision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inputs lead to significant computational and memory overhead, posing challenges for real-time deployment on hardware accelerators. While prior work attempts to reduce redundancy via token pruning or merging, these methods typically operate at coarse granularity and incur high runtime overhead due to global token-level operations. In this study, we propose Focus, a Streaming Concentration Architecture that efficiently accelerates VLM inference through progressive, fine-grained redundancy elimination. Focus introduces a multilevel concentration paradigm that hierarchically compresses vision-language inputs at three levels: (1) semantic-guided token pruning based on textual prompts, (2) spatial-temporal block-level…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Advanced Neural Network Applications · Advanced Image and Video Retrieval Techniques