Yanmin Qian
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Recognizing Multi-talker Speech with Permutation Invariant Training
2017 · arXiv (Cornell University)
In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) …
-
Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks
2018
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the …
-
Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we …
-
Self-Supervised Speaker Verification Using Dynamic Loss-Gate and Label Correction
2022 · arXiv (Cornell University)
For self-supervised speaker verification, the quality of pseudo labels decides the upper bound of the system due to the massive unreliable labels. In this work, we propose dynamic loss-gate and label correction (DLG-LC) to alleviate …
-
The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines
2022 · arXiv (Cornell University)
The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each …
-
Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
2022 · arXiv (Cornell University)
Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker …
-
Exploring Binary Classification Loss For Speaker Verification
2023 · arXiv (Cornell University)
The mismatch between close-set training and open-set testing usually leads to significant performance degradation for speaker verification task. For existing loss functions, metric learning-based objectives depend strongly on searching effective pairs which might hinder further …
-
Efficient Pruning for Large-Scale Seq2Seq Speech Models without Back-Propagation
2025
Large-scale Seq2Seq speech models like Whisper excel in speech recognition but are limited by their high computational demands, making them difficult to be deployed on resource-constrained devices. This paper introduces a novel and efficient pruning …
-
Target Speech Detection With Multimodal Prompts
2025 · IEEE Transactions on Audio Speech and Language Processing
Traditional speaker diarization seeks to detect “who spoke when” according to speaker characteristics. Extending to target speech detection, we detect “when target speech event occurs” according to the semantic characteristics of speech. We propose a …