Voxtral Realtime
Mistral-AI: Alexander H. Liu, Andy Ehrenberg, Andy Lo, Chen-Yo Sun, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Rohin Arora, Sanchit Gandhi, Sandeep Subramanian, Soham Ghosh, Srijan Mishra, Abhinav Rastogi

TL;DR
Voxtral Realtime is a streaming speech recognition model trained end-to-end for sub-second latency, achieving offline-quality transcription across 13 languages without chunking or sliding windows.
Contribution
It introduces a novel end-to-end streaming ASR architecture with explicit audio-text alignment and scalable multilingual pretraining.
Findings
Achieves offline-quality transcription at 480ms delay.
Performs on par with Whisper across 13 languages.
Released under Apache 2.0 license.
Abstract
We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adapt offline models through chunking or sliding windows, Voxtral Realtime is trained end-to-end for streaming, with explicit alignment between audio and text streams. Our architecture builds on the Delayed Streams Modeling framework, introducing a new causal audio encoder and Ada RMS-Norm for improved delay conditioning. We scale pretraining to a large-scale dataset spanning 13 languages. At a delay of 480ms, Voxtral Realtime achieves performance on par with Whisper, the most widely deployed offline transcription system. We release the model weights under the Apache 2.0 license.
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
- 🤗mistralai/Voxtral-Mini-4B-Realtime-2602model· 1.3M dl· ♡ 8571.3M dl♡ 857
- 🤗freddm/Voxtral-Mini-4B-Realtime-2602-GGUFmodel· 413 dl· ♡ 7413 dl♡ 7
- 🤗thedruid831/Voxtral-Mini-4B-Realtime-2602model· 12 dl12 dl
- 🤗RedHatAI/Voxtral-Mini-4B-Realtime-2602model· 52 dl52 dl
- 🤗tantk/Voxtral-4B-Realtime-VQFmodel· 4 dl4 dl
- 🤗onnx-community/Voxtral-Mini-4B-Realtime-2602-ONNXmodel· 2.8k dl· ♡ 222.8k dl♡ 22
- 🤗younghan-meta/Voxtral-Mini-4B-Realtime-2602-ExecuTorch-XNNPACKmodel· 10 dl10 dl
- 🤗younghan-meta/Voxtral-Mini-4B-Realtime-2602-ExecuTorch-Metalmodel· 27 dl27 dl
- 🤗Jacques976/Voxtral-Mini-4B-Realtime-2602model· 5 dl5 dl
- 🤗beaupi/Voxtral-Mini-4B-Realtime-2602-oQ8model· 11 dl11 dl
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
