ملف الباحث
Hosung Nam
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Text-to-speech with linear spectrogram prediction for quality and speed improvement
2021 · Phonetics and Speech Sciences
Most neural-network-based speech synthesis models utilize neural vocoders to convert mel-scaled spectrograms into high-quality, human-like voices. However, neural vocoders combined with mel-scaled spectrogram prediction models demand considerable computer memory and time during the training phase …
-
EARSHOT: A minimal neural network model of incremental human speech recognition
2018
Despite the “lack of invariance problem” (multiple acoustic patterns map to the same phoneme, and one acoustic pattern can map to different phonemes), humans experience phonetic constancy: we typically perceive what the speaker intended despite …