A Proximal Operator for Inducing 2:4-Sparsity

Jonas M K\"ubler; Yu-Xiang Wang; Shoham Sabach; Navid Ansari; Matth\"aus Kleindessner; Kailash Budhathoki; Volkan Cevher; George Karypis

arXiv:2501.18015·cs.LG·August 28, 2025

A Proximal Operator for Inducing 2:4-Sparsity

Jonas M K\"ubler, Yu-Xiang Wang, Shoham Sabach, Navid Ansari, Matth\"aus Kleindessner, Kailash Budhathoki, Volkan Cevher, George Karypis

PDF

Open Access

TL;DR

This paper introduces a novel regularizer and proximal operator to induce 2:4 sparsity in neural networks, improving pruning efficiency and accuracy, especially for large language models, by exploiting local feature correlations.

Contribution

We derive a proximal operator for 2:4 sparsity that enhances model pruning, achieving state-of-the-art results on large language models with minimal accuracy loss.

Findings

01

Improved pruning of models up to 13B parameters.

02

Matched state-of-the-art performance on 70B models.

03

Efficient solution for 2:4 sparsity regularization.

Abstract

Recent hardware advancements in AI Accelerators and GPUs allow to efficiently compute sparse matrix multiplications, especially when 2 out of 4 consecutive weights are set to zero. However, this so-called 2:4 sparsity usually comes at a decreased accuracy of the model. We derive a regularizer that exploits the local correlation of features to find better sparsity masks in trained models. We minimize the regularizer jointly with a local squared loss by deriving the proximal operator for which we show that it has an efficient solution in the 2:4-sparse case. After optimizing the mask, we use maskedgradient updates to further minimize the local squared loss. We illustrate our method on toy problems and apply it to pruning entire large language models up to 70B parameters. On models up to 13B we improve over previous state of the art algorithms, whilst on 70B models we match their…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsApproximation Theory and Sequence Spaces · Matrix Theory and Algorithms

MethodsPruning · Sparse Evolutionary Training