Efficient Ternary Weight Embedding Model: Bridging Scalability and   Performance

Jiayi Chen; Chen Wu; Shaoqun Zhang; Nan Li; Liangjie Zhang; Qi; Zhang

arXiv:2411.15438·cs.CL·November 26, 2024

Efficient Ternary Weight Embedding Model: Bridging Scalability and Performance

Jiayi Chen, Chen Wu, Shaoqun Zhang, Nan Li, Liangjie Zhang, Qi, Zhang

PDF

Open Access 1 Repo

TL;DR

This paper introduces a novel finetuning framework for ternary-weight embedding models that significantly reduces memory and computational costs while maintaining high performance, suitable for resource-constrained environments.

Contribution

It presents a self-taught knowledge distillation method for applying ternarization to pre-trained embedding models, enhancing efficiency without sacrificing effectiveness.

Findings

01

Ternary embedding models achieve low memory usage and low latency.

02

Combining ternary embeddings with ANN search improves accuracy and efficiency.

03

Extensive experiments validate the effectiveness across text and vision datasets.

Abstract

Embedding models have become essential tools in both natural language processing and computer vision, enabling efficient semantic search, recommendation, clustering, and more. However, the high memory and computational demands of full-precision embeddings pose challenges for deployment in resource-constrained environments, such as real-time recommendation systems. In this work, we propose a novel finetuning framework to ternary-weight embedding models, which reduces memory and computational overhead while maintaining high performance. To apply ternarization to pre-trained embedding models, we introduce self-taught knowledge distillation to finalize the ternary-weights of the linear layers. With extensive experiments on public text and vision datasets, we demonstrated that without sacrificing effectiveness, the ternarized model consumes low memory usage and has low latency in the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

dataparameters/Ternary-Embedding-Models
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Computing and Algorithms

MethodsKnowledge Distillation