CDLM: Consistency Diffusion Language Models For Faster Sampling

Minseo Kim; Chenfeng Xu; Coleman Hooper; Harman Singh; Ben Athiwaratkun; Ce Zhang; Kurt Keutzer; Amir Gholami

arXiv:2511.19269·cs.LG·February 23, 2026

CDLM: Consistency Diffusion Language Models For Faster Sampling

Minseo Kim, Chenfeng Xu, Coleman Hooper, Harman Singh, Ben Athiwaratkun, Ce Zhang, Kurt Keutzer, Amir Gholami

PDF

Open Access 2 Models

TL;DR

CDLM introduces a training-based approach to accelerate diffusion language models by reducing sampling steps and enabling KV caching, significantly lowering inference latency while maintaining accuracy.

Contribution

The paper proposes CDLM, a novel method that combines consistency modeling and block-wise causal attention to speed up diffusion language models.

Findings

01

Achieves 3.6x-14.5x lower latency in inference.

02

Maintains competitive accuracy on math and coding tasks.

03

Fully compatible with KV caching during inference.

Abstract

Diffusion Language Models (DLMs) offer a promising parallel generation paradigm but suffer from slow inference due to numerous refinement steps and the inability to use standard KV caching. We introduce CDLM (Consistency Diffusion Language Models), a training-based acceleration method that simultaneously tackles both bottlenecks. CDLM integrates consistency modeling to drastically reduce the number of required sampling steps by enabling multi-token finalization. Furthermore, we enforce a block-wise causal attention mask during fine-tuning, making the model fully compatible with KV caching. Experiments show CDLM achieves 3.6x-14.5x lower latency while maintaining competitive accuracy on math and coding tasks. The full training and evaluation code is available at https://github.com/SqueezeAILab/CDLM.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Machine Learning in Healthcare