conference-paper
Open access
SSR: Alignment-Aware Modality Connector for Speech Language Models
Research footprint
At a glance
- Citations
- 2
- References
- 0
- Comments
- 0
Paper overview
Abstract
Fusing speech into a pre-trained language model (SpeechLM) usually suffers from the inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality.We propose SSR-CONNECTOR (Segmented Speech Representation Connector) for better modality fusion.Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings.Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting.SSR-CONNECTOR outperforms existing mechanism for speechtext modality fusion, consistently achieving better speech understanding (e.g.
Record transparency
Publication details
- DOI
- 10.18653/v1/2025.iwslt-1.5
- OpenAlex
- W4412944344
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.