conference-paper وصول مفتوح

Subword and Crossword Units for CTC Acoustic Models

Research footprint

At a glance

الاستشهادات
3
المراجع
21
Comments
0
Paper overview

Abstract

This paper proposes a novel approach to create an unit set for CTC based speech recognition systems. By using Byte Pair Encoding we learn an unit set of an arbitrary size on a given training text. In contrast to using characters or words as units this allows us to find a good trade-off between the size of our unit set and the available training data. We evaluate both Crossword units, that may span multiple word, and Subword units. By combining this approach with decoding methods using a separate language model we are able to achieve state of the art results for grapheme based CTC systems.

Record transparency

Publication details

DOI
10.21437/interspeech.2018-2057
OpenAlex
W2779433219
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.