conference-paper

A long-duration Speech Semantic Recognition and Summarization Model for multi-speaker Conversations

Research footprint

At a glance

الاستشهادات
0
المراجع
17
Comments
0
Paper overview

Abstract

This paper aims to propose a comprehensive solution to address challenges in the field of audio processing for long-duration, multi-speaker conversational meetings. The identified issues include multi-speaker voice separation, difficulties in handling long temporal sequences, and the high redundancy in generating speech scene text. The proposed solution encompasses four modules: Firstly, addressing the issue of complex background noise interference in speech signals by introducing Deep Extractor for Music Sources with extra Components (Demucs) audio separation and denoising technology to effectively extract and suppress noise. Secondly, employing a combined approach of Convolutional Neural Network (CNN) and self-attention to extract multi-scale features for speaker identification in a multi-speaker environment. Thirdly, the Conformer model is trained using a large-scale Chinese speech database to tackle challenges in processing extended temporal sequences, ensuring the accurate transformation of speech information into textual representations. Finally, the latest Generative Pre-trained Transformer-3.5 (GPT-3.5) model undergoes fine-tuning using a Chinese text summarization dataset. This process enables the generation of text summaries, extracting key points from multi-speaker long-duration conversations. The effectiveness of the proposed algorithms is validated through speech scene experiments in various scenarios.

Record transparency

Publication details

DOI
10.23919/ccc63176.2024.10662739
OpenAlex
W4402570993
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.