Optimal Quantization for Matrix Multiplication

Or Ordentlich; Yury Polyanskiy

arXiv:2410.13780·cs.IT·October 16, 2025

Optimal Quantization for Matrix Multiplication

Or Ordentlich, Yury Polyanskiy

PDF

Open Access 2 Repos

TL;DR

This paper develops a universal quantization method for matrix multiplication that is theoretically optimal for Gaussian matrices, providing bounds, practical algorithms, and insights into rate-distortion behavior.

Contribution

It introduces a non-asymptotic lower bound and a universal lattice-based quantizer for matrix multiplication, achieving asymptotic optimality for Gaussian matrices and analyzing rate-distortion phase transitions.

Findings

01

Achieves asymptotic optimality for iid Gaussian matrices.

02

Provides a practical low-complexity quantizer with near-optimal performance.

03

Identifies a phase transition in rate-distortion at approximately 0.906 bits per entry.

Abstract

Recent work in machine learning community proposed multiple methods for performing lossy compression (quantization) of large matrices. This quantization is important for accelerating matrix multiplication (main component of large language models), which is often bottlenecked by the speed of loading these matrices from memory. Unlike classical vector quantization and rate-distortion theory, the goal of these new compression algorithms is to be able to approximate not the matrices themselves, but their matrix product. Specifically, given a pair of real matrices $A, B$ an encoder (compressor) is applied to each of them independently producing descriptions with $R$ bits per entry. These representations subsequently are used by the decoder to estimate matrix product $A^{⊤} B$ . In this work, we provide a non-asymptotic lower bound on the mean squared error of this approximation (as a function…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMedical Image Segmentation Techniques · Brain Tumor Detection and Classification · Mathematical Analysis and Transform Methods

MethodsSPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings