Direction is what you need: Improving Word Embedding Compression in   Large Language Models

Klaudia Ba{\l}azy; Mohammadreza Banaei; R\'emi Lebret; Jacek Tabor,; Karl Aberer

arXiv:2106.08181·cs.CL·August 4, 2021

Direction is what you need: Improving Word Embedding Compression in Large Language Models

Klaudia Ba{\l}azy, Mohammadreza Banaei, R\'emi Lebret, Jacek Tabor,, Karl Aberer

PDF

1 Repo

TL;DR

This paper introduces a novel autoencoder-based method focusing on embedding direction to effectively compress token embeddings in Transformer models, improving performance over traditional methods without additional pre-training.

Contribution

It proposes a task-agnostic embedding compression technique emphasizing directionality, outperforming SVD-based methods in language modeling and downstream tasks.

Findings

01

Outperforms SVD-based matrix factorization in perplexity

02

Achieves better results on SQuAD v1.1 and GLUE tasks

03

Does not require additional language model pre-training

Abstract

The adoption of Transformer-based models in natural language processing (NLP) has led to great success using a massive number of parameters. However, due to deployment constraints in edge devices, there has been a rising interest in the compression of these models to improve their inference time and memory footprint. This paper presents a novel loss objective to compress token embeddings in the Transformer-based models by leveraging an AutoEncoder architecture. More specifically, we emphasize the importance of the direction of compressed embeddings with respect to original uncompressed embeddings. The proposed method is task-agnostic and does not require further language modeling pre-training. Our method significantly outperforms the commonly used SVD-based matrix-factorization approach in terms of initial language model Perplexity. Moreover, we evaluate our proposed approach over SQuAD…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

MohammadrezaBanaei/orientation_based_embedding_compression
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.