Researcher profile

Kai Yu

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Agent-Aware Dropout DQN for Safe and Efficient On-line Dialogue Policy Learning

    2017

    Hand-crafted rules and reinforcement learning (RL) are two popular choices to obtain dialogue policy. The rule-based policy is often reliable within predefined scope but not self-adaptable, whereas RL is evolvable with data but often suffers …

  2. Ordering-Based Kalman Filter Selective Ensemble for Classification

    2020 · IEEE Access

    This paper investigates Kalman Filter-based Heuristic Ensemble (KFHE), which is a new perspective on multi-class ensemble classification with performance significantly better or at least as good as the state-of-the-art algorithms. We prove that the sample …

  3. Unsupervised Word-Level Prosody Tagging for Controllable Speech Synthesis

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Although word-level prosody modeling in neural text-to-speech (TTS) has been investigated in recent research for diverse speech synthesis, it is still challenging to control speech synthesis manually without a specific reference. This is largely due …

  4. Quantum K-nearest neighbor classification algorithm based on Hamming distance

    2021 · arXiv (Cornell University)

    K-nearest neighbor classification algorithm is one of the most basic algorithms in machine learning, which determines the sample's category by the similarity between samples. In this paper, we propose a quantum K-nearest neighbor classification algorithm …

  5. Reliable Federated Disentangling Network for Non-IID Domain Feature

    2023 · arXiv (Cornell University)

    Federated learning (FL), as an effective decentralized distributed learning approach, enables multiple institutions to jointly train a model without sharing their local data. However, the domain feature shift caused by different acquisition devices/clients substantially degrades …

  6. Exploring Schema Generalizability of Text-to-SQL

    2023

    Exploring the generalizability of a text-to-SQL parser is essential for a system to automatically adapt the real-world databases. Previous investigation works mostly focus on lexical diversity, including the influence of the synonym and perturbations in …

  7. StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

    2024

    While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of artistic works. In this paper, we introduce StoryTTS, a highly ETTS dataset …

  8. SciDFM: A Large Language Model with Mixture-of-Experts for Science

    2024 · arXiv (Cornell University)

    Recently, there has been a significant upsurge of interest in leveraging large language models (LLMs) to assist scientific discovery. However, most LLMs only focus on general science, while they lack domain-specific knowledge, such as chemical …

  9. Unified Pathological Speech Analysis with Prompt Tuning

    2024 · arXiv (Cornell University)

    Pathological speech analysis has been of interest in the detection of certain diseases like depression and Alzheimer's disease and attracts much interest from researchers. However, previous pathological speech analysis models are commonly designed for a …

  10. A multi-channel meter smart communication method based on low-voltage power line broadband carrier

    2025 · Journal of Physics Conference Series

    Abstract The conventional meter intelligent communication method establishes a communication network with an existing network, which is affected by background noise, impulse noise, and so on, and the communication quality is degraded. Therefore, a multi-channel …

  11. Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

    2025 · arXiv (Cornell University)

    RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the prediction network in RNN-T poses an overfitting issue, especially under challenging scenarios, …

  12. CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate

    2026

    Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. However, speech is inherently …

  13. Bidirectional LSTM-CRF Models for Sequence Tagging

    2015 · arXiv (Cornell University)

    At the moment, the vast majority of Portuguese archives with an online presence use a software solution to manage their finding aids: e.g. Digitarq or Archeevo. Most of these finding aids are written in natural …

  14. Towards Universal Dialogue State Tracking

    2018

    Dialogue state tracking is the core part of a spoken dialogue system. It estimates the beliefs of possible user's goals at every dialogue turn. However, for most current approaches, it's difficult to scale to large …