Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems

Haoyuan Yu; Yuxuan Chen; Minjie Cai

arXiv:2601.20230·cs.CL·January 30, 2026

Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems

Haoyuan Yu, Yuxuan Chen, Minjie Cai

PDF

Open Access

TL;DR

This paper introduces a novel framework for full-duplex dialogue systems that decomposes conversations into minimal units for independent processing, enhancing natural interaction capabilities.

Contribution

It proposes a semi-cascaded, train-free dialogue system using a multimodal large language model with auxiliary modules, improving full-duplex interaction performance.

Findings

01

Achieved second place on the Human-like Spoken Dialogue Systems Challenge.

02

Demonstrated effectiveness of unit-based decomposition in full-duplex dialogue.

03

Operates in a plug-and-play, train-free manner.

Abstract

Full-duplex voice interaction is crucial for natural human computer interaction. We present a framework that decomposes complex dialogue into minimal conversational units, enabling the system to process each unit independently and predict when to transit to the next. This framework is instantiated as a semi-cascaded full-duplex dialogue system built around a multimodal large language model, supported by auxiliary modules such as voice activity detection (VAD) and text-to-speech (TTS) synthesis. The resulting system operates in a train-free, plug-and-play manner. Experiments on the HumDial dataset demonstrate the effectiveness of our framework, which ranks second among all teams on the test set of the Human-like Spoken Dialogue Systems Challenge (Track 2: Full-Duplex Interaction). Code is available at the GitHub repository https://github.com/yu-haoyuan/fd-badcat.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and dialogue systems · Speech Recognition and Synthesis · Topic Modeling