conference-paper وصول مفتوح

SSR: Alignment-Aware Modality Connector for Speech Language Models

Research footprint

At a glance

الاستشهادات
2
المراجع
0
Comments
0
Paper overview

Abstract

Fusing speech into a pre-trained language model (SpeechLM) usually suffers from the inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality.We propose SSR-CONNECTOR (Segmented Speech Representation Connector) for better modality fusion.Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings.Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting.SSR-CONNECTOR outperforms existing mechanism for speechtext modality fusion, consistently achieving better speech understanding (e.g.

Record transparency

Publication details

DOI
10.18653/v1/2025.iwslt-1.5
OpenAlex
W4412944344
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.