Kai Yu
14 papers in the PaperMetrix corpus
Papers by this author
-
Agent-Aware Dropout DQN for Safe and Efficient On-line Dialogue Policy Learning
2017
Hand-crafted rules and reinforcement learning (RL) are two popular choices to obtain dialogue policy. The rule-based policy is often reliable within predefined scope but not self-adaptable, whereas RL is evolvable with data but often suffers …
-
Ordering-Based Kalman Filter Selective Ensemble for Classification
2020 · IEEE Access
This paper investigates Kalman Filter-based Heuristic Ensemble (KFHE), which is a new perspective on multi-class ensemble classification with performance significantly better or at least as good as the state-of-the-art algorithms. We prove that the sample …
-
Unsupervised Word-Level Prosody Tagging for Controllable Speech Synthesis
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Although word-level prosody modeling in neural text-to-speech (TTS) has been investigated in recent research for diverse speech synthesis, it is still challenging to control speech synthesis manually without a specific reference. This is largely due …
-
Quantum K-nearest neighbor classification algorithm based on Hamming distance
2021 · arXiv (Cornell University)
K-nearest neighbor classification algorithm is one of the most basic algorithms in machine learning, which determines the sample's category by the similarity between samples. In this paper, we propose a quantum K-nearest neighbor classification algorithm …
-
Reliable Federated Disentangling Network for Non-IID Domain Feature
2023 · arXiv (Cornell University)
Federated learning (FL), as an effective decentralized distributed learning approach, enables multiple institutions to jointly train a model without sharing their local data. However, the domain feature shift caused by different acquisition devices/clients substantially degrades …
-
Exploring Schema Generalizability of Text-to-SQL
2023
Exploring the generalizability of a text-to-SQL parser is essential for a system to automatically adapt the real-world databases. Previous investigation works mostly focus on lexical diversity, including the influence of the synonym and perturbations in …
-
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
2024
While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of artistic works. In this paper, we introduce StoryTTS, a highly ETTS dataset …
-
SciDFM: A Large Language Model with Mixture-of-Experts for Science
2024 · arXiv (Cornell University)
Recently, there has been a significant upsurge of interest in leveraging large language models (LLMs) to assist scientific discovery. However, most LLMs only focus on general science, while they lack domain-specific knowledge, such as chemical …
-
Unified Pathological Speech Analysis with Prompt Tuning
2024 · arXiv (Cornell University)
Pathological speech analysis has been of interest in the detection of certain diseases like depression and Alzheimer's disease and attracts much interest from researchers. However, previous pathological speech analysis models are commonly designed for a …
-
A multi-channel meter smart communication method based on low-voltage power line broadband carrier
2025 · Journal of Physics Conference Series
Abstract The conventional meter intelligent communication method establishes a communication network with an existing network, which is affected by background noise, impulse noise, and so on, and the communication quality is degraded. Therefore, a multi-channel …
-
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
2025 · arXiv (Cornell University)
RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the prediction network in RNN-T poses an overfitting issue, especially under challenging scenarios, …
-
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
2026
Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. However, speech is inherently …
-
Bidirectional LSTM-CRF Models for Sequence Tagging
2015 · arXiv (Cornell University)
At the moment, the vast majority of Portuguese archives with an online presence use a software solution to manage their finding aids: e.g. Digitarq or Archeevo. Most of these finding aids are written in natural …
-
Towards Universal Dialogue State Tracking
2018
Dialogue state tracking is the core part of a spoken dialogue system. It estimates the beliefs of possible user's goals at every dialogue turn. However, for most current approaches, it's difficult to scale to large …