Researcher profile

Mark Hasegawa‐Johnson

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks

    2018 · arXiv (Cornell University)

    Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions …

  2. Training Spoken Language Understanding Systems with Non-Parallel Speech and Text

    2020

    End-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate …

  3. Autosegmental Neural Nets: Should Phones and Tones be Synchronous or Asynchronous?

    2020

    Phones, the segmental units of the International Phonetic Alphabet (IPA), are used for lexical distinctions in most human languages; Tones, the suprasegmental units of the IPA, are used in perhaps 70%. Many previous studies have …

  4. A Theory of Unsupervised Speech Recognition

    2023 · arXiv (Cornell University)

    Unsupervised speech recognition (ASR-U) is the problem of learning automatic speech recognition (ASR) systems from unpaired speech-only and text-only corpora. While various algorithms exist to solve this problem, a theoretical framework is missing from studying …

  5. Unsupervised Speech Recognition with N-Skipgram and Positional Unigram Matching

    2023 · arXiv (Cornell University)

    Training unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the …

  6. LIMMITS’24: Multi-Speaker, Multi-Lingual Indic TTS with Voice Cloning

    2024

    The Multi-speaker, Multi-lingual Indic TTS with voice cloning (LIMMITS’24) challenge is organized as part of the ICASSP 2024 signal processing grand challenge. LIMMITS’24 aims at the development of voice cloning for multi-speaker, multi-lingual Text-to-Speech (TTS) …

  7. Streaming Recommender Systems

    2017

    The increasing popularity of real-world recommender systems produces data continuously and rapidly, and it becomes more realistic to study recommender systems under streaming scenarios. Data streams present distinct properties such as temporally ordered, continuous and …