preprint
وصول مفتوح
ESPnet: End-to-End Speech Processing Toolkit
Research footprint
At a glance
- الاستشهادات
- 74
- المراجع
- 36
- Comments
- 0
Paper overview
Abstract
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1804.00015
- OpenAlex
- W2795935804
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.