Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025

Dominik Mach\'a\v{c}ek; Peter Pol\'ak

arXiv:2506.17077·cs.CL·June 23, 2025

Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025

Dominik Mach\'a\v{c}ek, Peter Pol\'ak

PDF

Open Access 1 Video

TL;DR

This paper presents CUNI's simultaneous speech translation systems for IWSLT 2025, leveraging offline Whisper models, advanced policies, and context adaptation to improve translation quality across four language pairs.

Contribution

The paper introduces a comprehensive approach combining offline Whisper models with novel policies and context handling, achieving significant BLEU improvements over baselines.

Findings

01

2 BLEU point improvement on Czech-English

02

13-22 BLEU point improvements on English-German, Chinese, Japanese

03

Proposed new speech recognition latency measure

Abstract

This paper describes Charles University submission to the Simultaneous Speech Translation Task of the IWSLT 2025. We cover all four language pairs with a direct or cascade approach. The backbone of our systems is the offline Whisper speech model, which we use for both translation and transcription in simultaneous mode with the state-of-the-art simultaneous policy AlignAtt. We further improve the performance by prompting to inject in-domain terminology, and we accommodate context. Our cascaded systems further use EuroLLM for unbounded simultaneous translation. Compared to the Organizers' baseline, our systems improve by 2 BLEU points on Czech to English and 13-22 BLEU points on English to German, Chinese and Japanese on the development sets. Additionally, we also propose a new enhanced measure of speech recognition latency.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025· underline

Taxonomy

TopicsNatural Language Processing Techniques · Speech Recognition and Synthesis · Speech and dialogue systems