Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person
Lucas Rafael Stefanel Gris, Ricardo Marcacini, Arnaldo Candido, Junior, Edresson Casanova, Anderson Soares, Sandra Maria Alu\'isio

TL;DR
This study evaluates OpenAI's Whisper ASR model for punctuation prediction and topic modeling in Portuguese, highlighting its strengths and areas needing improvement in real-world applications involving life histories.
Contribution
First comprehensive assessment of Whisper's performance on Portuguese punctuation prediction and topic modeling in a real-world museum context.
Findings
Whisper achieves state-of-the-art punctuation prediction results.
Performance varies across different punctuation marks, with some needing improvement.
Effective for practical applications like topic modeling in virtual museums.
Abstract
Automatic speech recognition (ASR) systems play a key role in applications involving human-machine interactions. Despite their importance, ASR models for the Portuguese language proposed in the last decade have limitations in relation to the correct identification of punctuation marks in automatic transcriptions, which hinder the use of transcriptions by other systems, models, and even by humans. However, recently Whisper ASR was proposed by OpenAI, a general-purpose speech recognition model that has generated great expectations in dealing with such limitations. This chapter presents the first study on the performance of Whisper for punctuation prediction in the Portuguese language. We present an experimental evaluation considering both theoretical aspects involving pausing points (comma) and complete ideas (exclamation, question, and fullstop), as well as practical aspects involving…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsMultimodal Machine Learning Applications · Speech and dialogue systems · Speech Recognition and Synthesis
