Event-Priori-Based Vision-Language Model for Efficient Visual Understanding

Haotong Qin; Cheng Hu; Michele Magno

arXiv:2506.07627·cs.CV·June 10, 2025

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding

Haotong Qin, Cheng Hu, Michele Magno

PDF

Open Access

TL;DR

This paper introduces EP-VLM, a novel vision-language model that uses event-based motion priors to sparsify visual inputs, significantly reducing computational costs while maintaining high accuracy, enabling efficient deployment on edge devices.

Contribution

EP-VLM leverages motion priors from dynamic event vision to guide input sparsification and employs a position-preserving tokenization strategy, enhancing efficiency without sacrificing accuracy.

Findings

01

50% FLOPs reduction compared to baseline

02

Retains 98% of original accuracy on RealWorldQA

03

Demonstrates effective event-guided visual input sparsification

Abstract

Large Language Model (LLM)-based Vision-Language Models (VLMs) have substantially extended the boundaries of visual understanding capabilities. However, their high computational demands hinder deployment on resource-constrained edge devices. A key source of inefficiency stems from the VLM's need to process dense and redundant visual information. Visual inputs contain significant regions irrelevant to text semantics, rendering the associated computations ineffective for inference. This paper introduces a novel Event-Priori-Based Vision-Language Model, termed EP-VLM. Its core contribution is a novel mechanism leveraging motion priors derived from dynamic event vision to enhance VLM efficiency. Inspired by human visual cognition, EP-VLM first employs event data to guide the patch-wise sparsification of RGB visual inputs, progressively concentrating VLM computation on salient regions of the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Advanced Neural Network Applications · Generative Adversarial Networks and Image Synthesis