Researcher profile

Lantian Li

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Speaker segmentation using deep speaker vectors for fast speaker change scenarios

    2017

    A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a …

  2. Deep Factorization for Speech Signal

    2017 · arXiv (Cornell University)

    Speech signals are complex intermingling of various informative factors, and this information blending makes decoding any of the individual factors extremely difficult. A natural idea is to factorize each speech frame into independent factors, though …

  3. VAE-based Domain Adaptation for Speaker Verification

    2019 · arXiv (Cornell University)

    Deep speaker embedding has achieved satisfactory performance in speaker verification. By enforcing the neural model to discriminate the speakers in the training set, deep speaker embedding (called `x-vectors`) can be derived from the hidden layers. …

  4. Neural Discriminant Analysis for Deep Speaker Embedding

    2020 · arXiv (Cornell University)

    Probabilistic Linear Discriminant Analysis (PLDA) is a popular tool in open-set classification/verification tasks. However, the Gaussian assumption underlying PLDA prevents it from being applied to situations where the data is clearly non-Gaussian. In this paper, …

  5. A Glance is Enough: Extract Target Sentence By Looking at A keyword

    2023 · arXiv (Cornell University)

    This paper investigates the possibility of extracting a target sentence from multi-talker speech using only a keyword as input. For example, in social security applications, the keyword might be "help", and the goal is to …

  6. Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions

    2024 · arXiv (Cornell University)

    Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due …