Researcher profile

Haizhou Li

11 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. On statistical machine translation method for lexicon refinement in speech recognition

    2015

    In low resource Automatic Speech Recognition (ASR), one usually resorts to the Statistical Machine Translation (SMT) technique to learn transform rules to refine grapheme lexicon. To do this, we face two challenges. One is to …

  2. Spoofing detection under noisy conditions: a preliminary investigation and an initial database

    2016 · arXiv (Cornell University)

    Spoofing detection for automatic speaker verification (ASV), which is to discriminate between live speech and attacks, has received increasing attentions recently. However, all the previous studies have been done on the clean data without significant …

  3. A Modularized Neural Network with Language-Specific Output Layers for Cross-lingual Voice Conversion

    2019 · arXiv (Cornell University)

    This paper presents a cross-lingual voice conversion framework that adopts a modularized neural network. The modularized neural network has a common input structure that is shared for both languages, and two separate output modules, one …

  4. HLT-NUS SUBMISSION FOR 2020 NIST Conversational Telephone Speech SRE

    2021 · arXiv (Cornell University)

    This work provides a brief description of Human Language Technology (HLT) Laboratory, National University of Singapore (NUS) system submission for 2020 NIST conversational telephone speech (CTS) speaker recognition evaluation (SRE). The challenge focuses on evaluation …

  5. LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT

    2022 · arXiv (Cornell University)

    Self-supervised speech representation learning has shown promising results in various speech processing tasks. However, the pre-trained models, e.g., HuBERT, are storage-intensive Transformers, limiting their scope of applications under low-resource settings. To this end, we propose …

  6. Just Rank: Rethinking Evaluation with Word and Sentence Similarities

    2022 · arXiv (Cornell University)

    Word and sentence embeddings are useful feature representations in natural language processing. However, intrinsic evaluation for embeddings lags far behind, and there has been no significant update since the past decade. Word and sentence similarity …

  7. Relational Sentence Embedding for Flexible Semantic Matching

    2023

    We present Relational Sentence Embedding (RSE), a new paradigm to further discover the potential of sentence embeddings.Prior work mainly models the similarity between sentences based on their embedding distance.Because of the complex semantic meanings conveyed, …

  8. RefXVC: Cross-Lingual Voice Conversion With Enhanced Reference Leveraging

    2024 · IEEE/ACM Transactions on Audio Speech and Language Processing

    This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker identity, which does …

  9. Target Speech Detection With Multimodal Prompts

    2025 · IEEE Transactions on Audio Speech and Language Processing

    Traditional speaker diarization seeks to detect “who spoke when” according to speaker characteristics. Extending to target speech detection, we detect “when target speech event occurs” according to the semantic characteristics of speech. We propose a …

  10. Utilizing Contextual Clues and Role Correlations for Enhancing Document-Level Event Argument Extraction

    2025 · IEEE Transactions on Audio Speech and Language Processing

    Document-level event argument extraction is a crucial yet challenging task within the field of information extraction. Current mainstream approaches primarily focus on the information interaction between event triggers and their arguments, facing two limitations: insufficient …

  11. Adequacy–Fluency Metrics: Evaluating MT in the Continuous Space Model Framework

    2015 · IEEE/ACM Transactions on Audio Speech and Language Processing

    This work extends and evaluates a two-dimensional automatic evaluation metric for machine translation, which is designed to operate at the sentence level. The metric is based on the concepts of adequacy and fluency, aiming at …