P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Yuzong Chen; Chao Fang; Xilai Dai; Yuheng Wu; Thierry Tambe; Marian Verhelst; Mohamed S. Abdelfattah

arXiv:2511.06838·cs.AR·May 5, 2026

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Yuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu, Thierry Tambe, Marian Verhelst, Mohamed S. Abdelfattah

PDF

1 Repo

TL;DR

P3-LLM is an innovative NPU-PIM accelerator that employs hybrid numerical formats and operator fusion to enhance edge LLM inference efficiency, accuracy, and speed.

Contribution

It introduces a flexible mixed-precision quantization scheme and a low-precision PIM architecture co-designed for improved LLM inference performance.

Findings

01

Achieves up to 4.9x speedup over state-of-the-art accelerators.

02

Maintains higher accuracy than existing quantization algorithms.

03

Demonstrates effective operator fusion to reduce runtime dequantization overhead.

Abstract

The substantial memory bandwidth and computational demands of large language models (LLMs) present critical challenges for efficient inference. To tackle this, the literature has explored heterogeneous systems that combine neural processing units (NPUs) with DRAM-based processing-in-memory (PIM) for LLM acceleration. However, the high-precision PIM compute units incur significant area and power overhead in DRAM technology, limiting the effective computation throughput. In this paper, we introduce P3-LLM, a novel NPU-PIM integrated accelerator for edge LLM inference. Our approach is threefold: First, we propose a flexible mixed-precision quantization scheme, which leverages hybrid numerical formats to quantize different LLM operands with high compression efficiency and minimal accuracy loss. Second, we architect an efficient PIM accelerator for P3-LLM, featuring enhanced compute units to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

yc2367/P3-LLM
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.