Researcher profile

Lei Xie

16 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Dual sub-swarm interaction QPSO algorithm based on different correlation coefficients

    2017 · Automatika

    A novel quantum-behaved particle swarm optimization (QPSO) algorithm, the dual sub-swarm interaction QPSO algorithm based on different correlation coefficients (DCC-QPSO), is proposed by constructing master-slave sub-swarms with different potential well centres. In the novel algorithm, …

  2. An Asynchronous WFST-Based Decoder For Automatic Speech Recognition

    2021 · arXiv (Cornell University)

    We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which …

  3. A Location Privacy Preservation Method Based on Dummy Locations in Internet of Vehicles

    2021 · Applied Sciences

    During the procedure, a location-based service (LBS) query, the real location provided by the vehicle user may results in the disclosure of vehicle location privacy. Moreover, the point of interest retrieval service requires high accuracy …

  4. Triangle Search Optimization Algorithm for Single-Objective Bound-Constrained Numerical Optimization

    2020 · 2020 5th International Conference on Mechanical, Control and Computer Engineering (ICMCCE)

    Real-parameter optimization has been a focus of the last decade. An inspired algorithm, triangle search optimization (TSO), is proposed for the Congress on Evolutionary Computation (CEC) 2020 competition. In this paper, the TSO algorithm is …

  5. WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit

    2021 · arXiv (Cornell University)

    In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify streaming and non-streaming end-to-end (E2E) speech recognition …

  6. Improving Robustness of One-Shot Voice Conversion with Deep Discriminative Speaker Encoder

    2021

    One-shot voice conversion has received significant attention since only one utterance from source speaker and target speaker respectively is required.Moreover, source speaker and target speaker do not need to be seen during training.However, available one-shot …

  7. Auto-KWS 2021 Challenge: Task, Datasets, and Baselines

    2021

    Auto-KWS 2021 challenge calls for automated machine learning (AutoML) solutions to automate the process of applying machine learning to a customized keyword spotting task.Compared with other keyword spotting tasks, Auto-KWS challenge has the following three …

  8. Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis

    2022 · Interspeech 2022

    Cross-speaker emotion transfer speech synthesis aims to synthesize emotional speech for a target speaker by transferring the emotion from reference speech recorded by another (source) speaker.In this task, extracting speaker-independent emotion embedding from reference speech …

  9. The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines

    2022 · arXiv (Cornell University)

    The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each …

  10. TSUP Speaker Diarization System for Conversational Short-phrase Speaker Diarization Challenge

    2022 · arXiv (Cornell University)

    This paper describes the TSUP team's submission to the ISCSLP 2022 conversational short-phrase speaker diarization (CSSD) challenge which particularly focuses on short-phrase conversations with a new evaluation metric called conversational diarization error rate (CDER). In …

  11. A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings

    2022 · arXiv (Cornell University)

    In this paper, we conduct a comparative study on speaker-attributed automatic speech recognition (SA-ASR) in the multi-party meeting scenario, a topic with increasing attention in meeting rich transcription. Specifically, three approaches are evaluated in this …

  12. ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge

    2024

    To promote speech processing and recognition research in driving scenarios, we build on the success of the Intelligent Cockpit Speech Recognition Challenge (ICSRC) held at ISCSLP 2022 and launch the ICASSP 2024 In-Car Multi-Channel Automatic …

  13. Learning Multi-view Anomaly Detection with Efficient Adaptive Selection

    2024 · arXiv (Cornell University)

    This study explores the recently proposed and challenging multi-view Anomaly Detection (AD) task. Single-view tasks will encounter blind spots from other perspectives, resulting in inaccuracies in sample-level prediction. Therefore, we introduce the Multi-View Anomaly Detection …

  14. HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models

    2024 · arXiv (Cornell University)

    Recent advancements in integrating Large Language Models (LLM) with automatic speech recognition (ASR) have performed remarkably in general domains. While supervised fine-tuning (SFT) of all model parameters is often employed to adapt pre-trained LLM-based ASR …

  15. Real-Time, Crowdsourcing-Enhanced Forecasting of Building Functionality During Urban Floods

    2026 · Computer-Aided Civil and Infrastructure Engineering

    Urban flood emergency response increasingly relies on infrastructure impact forecasts rather than hazard variables alone. However, real-time predictions are unreliable due to biased rainfall, incomplete flood knowledge, and sparse observations. Conventional open-loop forecasting propagates impacts …

  16. Investigating End-to-end Speech Recognition for Mandarin-english Code-switching

    2019

    Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, …