preprint وصول مفتوح

Improving Transformer-based Speech Recognition Using Unsupervised Pre-training

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
103
المراجع
30
Comments
0
Paper overview

Abstract

Speech recognition technologies are gaining enormous popularity in various industrial applications. However, building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect. To tackle this problem, an unsupervised pre-training method called Masked Predictive Coding is proposed, which can be applied for unsupervised pre-training with Transformer based model. Experiments on HKUST show that using the same training data, we can achieve CER 23.3%, exceeding the best end-to-end model by over 0.2% absolute CER. With more pre-training data, we can further reduce the CER to 21.0%, or a 11.8% relative CER reduction over baseline.

Record transparency

Publication details

DOI
10.48550/arxiv.1910.09932
OpenAlex
W2981991061
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.