Takaaki Hori
6 papers in the PaperMetrix corpus
Papers by this author
-
Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling
2018 · arXiv (Cornell University)
Sequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech research. The approach benefits by performing model training without using lexicon and alignments. However, this poses a new problem of requiring more …
-
Semi-Supervised Speech Recognition Via Graph-Based Temporal Classification
2021
Semi-supervised learning has demonstrated promising results in automatic speech recognition (ASR) by self-training using a seed ASR model with pseudo-labels generated for unlabeled data. The effectiveness of this approach largely relies on the pseudo-label accuracy, …
-
Joint CTC-attention based end-to-end speech recognition using multi-task learning
2017
Recently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping between variable-length input …
-
ESPnet: End-to-End Speech Processing Toolkit
2018 · arXiv (Cornell University)
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a …
-
Cycle-consistency Training for End-to-end Speech Recognition
2019
This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, …
-
A Comparative Study on Transformer vs RNN in Speech Applications
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …