CASPER: A Large Scale Spontaneous Speech Dataset

Cihan Xiao; Ruixing Liang; Xiangyu Zhang; Mehmet Emre Tiryaki; Veronica Bae; Lavanya Shankar; Rong Yang; Ethan Poon; Emmanuel Dupoux; Sanjeev Khudanpur; Leibny Paola Garcia Perera

arXiv:2506.00267·cs.CL·June 12, 2025

CASPER: A Large Scale Spontaneous Speech Dataset

Cihan Xiao, Ruixing Liang, Xiangyu Zhang, Mehmet Emre Tiryaki, Veronica Bae, Lavanya Shankar, Rong Yang, Ethan Poon, Emmanuel Dupoux, Sanjeev Khudanpur, Leibny Paola Garcia Perera

PDF

Open Access

TL;DR

This paper introduces CASPER, a large-scale dataset of over 100 hours of spontaneous speech collected through a novel dialogue elicitation pipeline, aiming to support speech processing research with more natural conversational data.

Contribution

It presents a new methodology for collecting high-quality spontaneous speech data and releases a substantial dataset to address the scarcity of natural dialogue resources.

Findings

01

Over 100 hours of spontaneous speech data collected

02

A reproducible pipeline for natural dialogue elicitation

03

Foundation for future expansion of spontaneous speech datasets

Abstract

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and dialogue systems · AI in Service Interactions · Speech Recognition and Synthesis