conference-paper Open access

CMU’s IWSLT 2025 Simultaneous Speech Translation System

Research footprint

At a glance

Citations
0
References
1
Comments
0
Paper overview

Abstract

This paper presents CMU's submission to the IWSLT 2025 Simultaneous Speech Translation (SST) task for translating unsegmented English speech into Chinese and German text in a streaming manner.Our end-to-end speechto-text system integrates a chunkwise causal Wav2Vec 2.0 speech encoder, an adapter, and the Qwen2.5-7B-Instructas the decoder.We use a two-stage simultaneous training procedure on robust speech segments curated from LibriSpeech, CommonVoice, and VoxPopuli datasets, utilizing standard cross-entropy loss.Our model supports adjustable latency through a configurable latency multiplier.Experimental results demonstrate that our system achieves 44.3 BLEU for English-to-Chinese and 25.1 BLEU for English-to-German translations on the ACL60/60 development set, with computation-aware latencies of 2.7 seconds and 2.3 seconds, and theoretical latencies of 2.2 and 1.7 seconds, respectively.

Record transparency

Publication details

DOI
10.18653/v1/2025.iwslt-1.31
OpenAlex
W4412944375
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.