Researcher profile

Helen Meng

17 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech

    2018

    For mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct …

  2. Recurrent Neural Network Language Model Training Using Natural Gradient

    2019

    Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, …

  3. Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention Networks

    2020

    Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet.In this paper, we present an XLNet-like pretraining scheme "Speech-XLNet" to learn speech representations with self-attention …

  4. Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization

    2021

    Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a promising method by using the limited dysarthric speech to fine-tune …

  5. Age-Invariant Speaker Embedding for Diarization of Cognitive Assessments

    2021

    This paper investigates an age-invariant speaker embedding approach to speaker diarization, which is an essential step towards the automatic cognitive assessments from speech. Studies have shown that incorporating speaker traits (e.g., age, gender, etc.) can …

  6. Controllable Emphatic Speech Synthesis based on Forward Attention for Expressive Speech Synthesis

    2021

    In speech interaction scenarios, speech emphasis is essential for expressing the underlying intention and attitude. Recently, end-to-end emphatic speech synthesis greatly improves the naturalness of synthetic speech, but also brings new problems: 1) lack of …

  7. Channel-wise Gated Res2Net: Towards Robust Detection of Synthetic Speech Attacks

    2021 · arXiv (Cornell University)

    Existing approaches for anti-spoofing in automatic speaker verification (ASV) still lack generalizability to unseen attacks. The Res2Net approach designs a residual-like connection between feature groups within one block, which increases the possible receptive fields and …

  8. Adversarially Learning Disentangled Speech Representations for Robust Multi-Factor Voice Conversion

    2021

    Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC).Conventional speech representation learning methods in VC only factorize speech as speaker and content, lacking controllability on other …

  9. A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

    2022 · arXiv (Cornell University)

    We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple …

  10. Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling

    2021 · arXiv (Cornell University)

    Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual …

  11. A Multi-Scale Time-Frequency Spectrogram Discriminator for GAN-based Non-Autoregressive TTS

    2022 · Interspeech 2022

    The generative adversarial network (GAN) has shown its outstanding capability in improving Non-Autoregressive TTS (NAR-TTS) by adversarially training it with an extra model that discriminates between the real and the generated speech.To maximize the benefits …

  12. An End-to-end Chinese Text Normalization Model based on Rule-guided Flat-Lattice Transformer

    2022 · arXiv (Cornell University)

    Text normalization, defined as a procedure transforming non standard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence …

  13. Towards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

    2022 · arXiv (Cornell University)

    Previous works on expressive speech synthesis focus on modelling the mono-scale style embedding from the current sentence or context, but the multi-scale nature of speaking style in human speech is neglected. In this paper, we …

  14. Context-aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis

    2023 · arXiv (Cornell University)

    Recent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in audiobooks. In this paper, …

  15. Enhancing the vocal range of single-speaker singing voice synthesis with melody-unsupervised pre-training

    2023 · arXiv (Cornell University)

    The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised …

  16. Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models

    2024 · arXiv (Cornell University)

    Self-supervised (SSL) models have shown great performance in various downstream tasks. However, they are typically developed for limited languages, and may encounter new languages in real-world. Developing a SSL model for each new language is …

  17. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

    2024 · arXiv (Cornell University)

    Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, …