conference-paper

SoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation

  • Interspeech 2022
Research footprint

At a glance

الاستشهادات
12
المراجع
19
Comments
0
Paper overview

Abstract

End-to-end speech synthesis models directly convert the input characters into an audio representation (e.g., spectrograms).Despite their impressive performance, such models have difficulty disambiguating the pronunciations of identically spelled words.To mitigate this issue, a separate Grapheme-to-Phoneme (G2P) model can be employed to convert the characters into phonemes before synthesizing the audio.This paper proposes SoundChoice, a novel G2P architecture that processes entire sentences rather than operating at the word level.The proposed architecture takes advantage of a weighted homograph loss (that improves disambiguation), exploits curriculum learning (that gradually switches from wordlevel to sentence-level G2P), and integrates word embeddings from BERT (for further performance improvement).Moreover, the model inherits the best practices in speech recognition, including multi-task learning with Connectionist Temporal Classification (CTC) and beam search with an embedded language model.As a result, SoundChoice achieves a Phoneme Error Rate (PER) of 2.65% on whole-sentence transcription using data from LibriSpeech and Wikipedia.

Record transparency

Publication details

DOI
10.21437/interspeech.2022-11066
OpenAlex
W4297841830
Document type
conference-paper
Language
EN
Source
Interspeech 2022
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.