Nanxin Chen
3 papers in the PaperMetrix corpus
Papers by this author
-
Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings
2019 · arXiv (Cornell University)
While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for …
-
ESPnet: End-to-End Speech Processing Toolkit
2018 · arXiv (Cornell University)
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a …
-
A Comparative Study on Transformer vs RNN in Speech Applications
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …