conference-paper Open access

SSR: Alignment-Aware Modality Connector for Speech Language Models

Research footprint

At a glance

Citations
2
References
0
Comments
0
Paper overview

Öz

Fusing speech into a pre-trained language model (SpeechLM) usually suffers from the inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality.We propose SSR-CONNECTOR (Segmented Speech Representation Connector) for better modality fusion.Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings.Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting.SSR-CONNECTOR outperforms existing mechanism for speechtext modality fusion, consistently achieving better speech understanding (e.g.

Record transparency

Publication details

DOI
10.18653/v1/2025.iwslt-1.5
OpenAlex
W4412944344
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.