Tomoki Hayashi
10 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Non-Parallel Voice Conversion System With WaveNet Vocoder and Collapsed Speech Suppression
2020 · IEEE Access
In this paper, we integrate a simple non-parallel voice conversion (VC) system with a WaveNet (WN) vocoder and a proposed collapsed speech suppression technique. The effectiveness of WN as a vocoder for generating high-fidelity speech …
-
ESPnet-ST: All-in-One Speech Translation Toolkit
2020 · arXiv (Cornell University)
We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly implements automatic …
-
Quasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech Generation
2020
In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) speech generation model, whose generative speed …
-
Non-Autoregressive Sequence-To-Sequence Voice Conversion
2021
This paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the FastSpeech2 model for …
-
On Prosody Modeling for ASR+TTS Based Voice Conversion
2021 · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to transcribe the source speech into the underlying …
-
A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
2022 · arXiv (Cornell University)
We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attractive owing to their potential to replace expensive supervised representations such as phonetic …
-
ESPnet: End-to-End Speech Processing Toolkit
2018 · arXiv (Cornell University)
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a …
-
Cycle-consistency Training for End-to-end Speech Recognition
2019
This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, …
-
A Comparative Study on Transformer vs RNN in Speech Applications
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …
-
Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
2020
This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and …