article

Target Speech Detection With Multimodal Prompts

  • IEEE Transactions on Audio Speech and Language Processing
Research footprint

At a glance

Citations
1
References
60
Comments
0
Paper overview

Abstract

Traditional speaker diarization seeks to detect “who spoke when” according to speaker characteristics. Extending to target speech detection, we detect “when target speech event occurs” according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Detection (MMTSD) framework, which accommodates diverse and multimodal prompts to specify target speech events in a flexible and userfriendly manner, including semantic language description, preenrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multimodal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.

Record transparency

Publication details

DOI
10.1109/taslpro.2025.3579304
OpenAlex
W4411232388
Document type
article
Language
EN
Source
IEEE Transactions on Audio Speech and Language Processing
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.