Language Models as Zero-shot Lossless Gradient Compressors: Towards   General Neural Parameter Prior Models

Hui-Po Wang; Mario Fritz

arXiv:2409.17836·cs.LG·January 23, 2025

Language Models as Zero-shot Lossless Gradient Compressors: Towards General Neural Parameter Prior Models

Hui-Po Wang, Mario Fritz

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces LM-GC, a novel method using large language models as zero-shot gradient priors for lossless compression, significantly improving compression efficiency and demonstrating potential for neural network gradient modeling.

Contribution

The paper presents LM-GC, a new approach that leverages LLMs with arithmetic coding to enhance gradient compression, a novel application of language models in this domain.

Findings

01

Token efficiency increased by up to 38 times.

02

Achieved 10% to 17.2% better compression rates than existing methods.

03

Demonstrated compatibility with lossy compression techniques.

Abstract

Despite the widespread use of statistical prior models in various fields, such models for neural network gradients have long been overlooked. The inherent challenge stems from their high-dimensional structures and complex interdependencies, which complicate effective modeling. In this work, we demonstrate the potential of large language models (LLMs) to act as gradient priors in a zero-shot setting. We examine the property by considering lossless gradient compression -- a critical application in distributed learning -- that depends heavily on precise probability modeling. To achieve this, we introduce LM-GC, a novel method that integrates LLMs with arithmetic coding. Our technique converts plain gradients into text-like formats, enhancing token efficiency by up to 38 times compared to their plain representations. We ensure that this data conversion maintains a close alignment with the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

hui-po-wang/LM-GC
pytorchOfficial

Videos

Language Models as Zero-shot Lossless Gradient Compressors: Towards General Neural Parameter Prior Models· slideslive

Taxonomy

TopicsSpeech Recognition and Synthesis