Nimbus: Secure and Efficient Two-Party Inference for Transformers

Zhengyi Li; Kang Yang; Jin Tan; Wen-jie Lu; Haoqi Wu; Xiao Wang; Yu; Yu; Derun Zhao; Yancheng Zheng; Minyi Guo; Jingwen Leng

arXiv:2411.15707·cs.CR·November 26, 2024

Nimbus: Secure and Efficient Two-Party Inference for Transformers

Zhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu, Haoqi Wu, Xiao Wang, Yu, Yu, Derun Zhao, Yancheng Zheng, Minyi Guo, Jingwen Leng

PDF

Open Access 1 Repo 1 Video

TL;DR

Nimbus introduces a secure, efficient two-party inference framework for Transformer models, significantly improving performance and maintaining high accuracy by optimizing matrix multiplication and non-linear layer computations.

Contribution

The paper proposes novel 2PC paradigms and polynomial approximations tailored for Transformers, enhancing efficiency and accuracy in privacy-preserving inference.

Findings

01

Achieves 2.9x to 12.5x faster matrix multiplication in linear layers.

02

Improves polynomial approximation performance by 2.9x to 4.0x with minimal accuracy loss.

03

Enhances end-to-end BERT inference speed by 2.7x to 4.7x.

Abstract

Transformer models have gained significant attention due to their power in machine learning tasks. Their extensive deployment has raised concerns about the potential leakage of sensitive information during inference. However, when being applied to Transformers, existing approaches based on secure two-party computation (2PC) bring about efficiency limitations in two folds: (1) resource-intensive matrix multiplications in linear layers, and (2) complex non-linear activation functions like $GELU$ and $Softmax$ . This work presents a new two-party inference framework $Nimbus$ for Transformer models. For the linear layer, we propose a new 2PC paradigm along with an encoding approach to securely compute matrix multiplications based on an outer-product insight, which achieves $2.9 \times \sim 12.5 \times$ performance improvements compared to the state-of-the-art (SOTA)…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

secretflow/spu
tfOfficial

Videos

Nimbus: Secure and Efficient Two-Party Inference for Transformers· slideslive

Taxonomy

TopicsCryptography and Data Security

MethodsAttention Is All You Need · Dense Connections · Label Smoothing · Dropout · Linear Layer · Layer Normalization · Byte Pair Encoding · Adam · Residual Connection · Softmax