Sei Ueno
3 papers in the PaperMetrix corpus
Papers by this author
-
End-to-end speech-to-dialog-act recognition
2020 · arXiv (Cornell University)
Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by …
-
Synthesizing waveform sequence-to-sequence to augment training data for sequence-to-sequence speech recognition
2021 · Nippon Onkyo Gakkaishi/Acoustical science and technology/Nihon Onkyo Gakkaishi
Sequence-to-sequence (seq2seq) automatic speech recognition (ASR) recently achieves state-of-the-art performance with fast decoding and a simple architecture. On the other hand, it requires a large amount of training data and cannot use text-only data for …
-
Refining Synthesized Speech Using Speaker Information and Phone Masking for Data Augmentation of Speech Recognition
2024 · IEEE/ACM Transactions on Audio Speech and Language Processing
While end-to-end automatic speech recognition (ASR) has shown impressive performance, it requires a huge amount of speech and transcription data. The conversion of domain-matched text to speech (TTS) has been investigated as one approach to …