ملف الباحث

Tomoki Hayashi

10 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Non-Parallel Voice Conversion System With WaveNet Vocoder and Collapsed Speech Suppression

    2020 · IEEE Access

    In this paper, we integrate a simple non-parallel voice conversion (VC) system with a WaveNet (WN) vocoder and a proposed collapsed speech suppression technique. The effectiveness of WN as a vocoder for generating high-fidelity speech …

  2. ESPnet-ST: All-in-One Speech Translation Toolkit

    2020 · arXiv (Cornell University)

    We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly implements automatic …

  3. Quasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech Generation

    2020

    In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) speech generation model, whose generative speed …

  4. Non-Autoregressive Sequence-To-Sequence Voice Conversion

    2021

    This paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the FastSpeech2 model for …

  5. On Prosody Modeling for ASR+TTS Based Voice Conversion

    2021 · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to transcribe the source speech into the underlying …

  6. A Comparative Study of Self-supervised Speech Representation Based Voice Conversion

    2022 · arXiv (Cornell University)

    We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attractive owing to their potential to replace expensive supervised representations such as phonetic …

  7. ESPnet: End-to-End Speech Processing Toolkit

    2018 · arXiv (Cornell University)

    This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a …

  8. Cycle-consistency Training for End-to-end Speech Recognition

    2019

    This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, …

  9. A Comparative Study on Transformer vs RNN in Speech Applications

    2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …

  10. Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

    2020

    This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and …