Jaň Černocký
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition
2021
Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the …
-
Probing Self-supervised Learning Models with Target Speech Extraction
2024 · arXiv (Cornell University)
Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-talker scenarios, such as extracting a target speaker in a mixture, is yet to …
-
Target Speech Extraction with Pre-Trained Self-Supervised Learning Models
2024
Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target …
-
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
2026
We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned Whisper (DiCoW) encoder to extract target-speaker embeddings, which are concatenated …