V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Tien-Yu Chi; Hung-Yueh Chiang; Chi-Chih Chang; Ning-Chi Huang,; Kai-Chiang Wu

arXiv:2412.16602·cs.CV·December 24, 2024

V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Tien-Yu Chi, Hung-Yueh Chiang, Chi-Chih Chang, Ning-Chi Huang,, Kai-Chiang Wu

PDF

Open Access

TL;DR

VMeanba is a novel, training-free compression technique for vision state space models that reduces computational complexity by averaging across channels, enabling faster image processing with minimal accuracy loss.

Contribution

Introduces VMeanba, a channel-averaging method that simplifies SSMs, improving efficiency without retraining, and extends their application to high-resolution vision tasks.

Findings

01

Achieves up to 1.12x speedup in image tasks

02

Maintains less than 3% accuracy loss

03

Effective with 40% unstructured pruning

Abstract

Vision transformers dominate image processing tasks due to their superior performance. However, the quadratic complexity of self-attention limits the scalability of these systems and their deployment on resource-constrained devices. State Space Models (SSMs) have emerged as a solution by introducing a linear recurrence mechanism, which reduces the complexity of sequence modeling from quadratic to linear. Recently, SSMs have been extended to high-resolution vision tasks. Nonetheless, the linear recurrence mechanism struggles to fully utilize matrix multiplication units on modern hardware, resulting in a computational bottleneck. We address this issue by introducing \textit{VMeanba}, a training-free compression method that eliminates the channel dimension in SSMs using mean operations. Our key observation is that the output activations of SSM blocks exhibit low variances across channels.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Visualization and Analytics