DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Shansan Gong; Mukai Li; Jiangtao Feng; Zhiyong Wu and; Lingpeng Kong

arXiv:2210.08933·cs.CL·February 15, 2023·94 cites

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu and, Lingpeng Kong

PDF

Open Access 1 Repo 1 Video

TL;DR

DiffuSeq introduces a diffusion model tailored for sequence-to-sequence text generation, achieving high quality and diversity, and demonstrating potential as an alternative to traditional autoregressive models.

Contribution

This paper presents the first diffusion-based model for Seq2Seq text generation, bridging diffusion models with natural language processing.

Findings

01

DiffuSeq performs comparably or better than six baselines, including state-of-the-art models.

02

DiffuSeq exhibits high diversity in generated outputs.

03

Theoretical analysis links DiffuSeq with autoregressive models.

Abstract

Recently, diffusion models have emerged as a new paradigm for generative models. Despite the success in domains using continuous signals such as vision and audio, adapting diffusion models to natural language is under-explored due to the discrete nature of texts, especially for conditional generation. We tackle this challenge by proposing DiffuSeq: a diffusion model designed for sequence-to-sequence (Seq2Seq) text generation tasks. Upon extensive evaluation over a wide range of Seq2Seq tasks, we find DiffuSeq achieving comparable or even better performance than six established baselines, including a state-of-the-art model that is based on pre-trained language models. Apart from quality, an intriguing property of DiffuSeq is its high diversity during generation, which is desired in many Seq2Seq tasks. We further include a theoretical analysis revealing the connection between DiffuSeq and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Shark-NLP/DiffuSeq
pytorchOfficial

Videos

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models· slideslive

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Music and Audio Processing

MethodsTanh Activation · Sigmoid Activation · Long Short-Term Memory · Sequence to Sequence · Diffusion