Dong Yu
16 papers in the PaperMetrix corpus
Papers by this author
-
Recognizing Multi-talker Speech with Permutation Invariant Training
2017 · arXiv (Cornell University)
In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) …
-
BLCU_NLP at SemEval-2018 Task 12: An Ensemble Model for Argument Reasoning Based on Hierarchical Attention
2018
The argument comprehension reasoning task aims to reconstruct and analyze the argument reasoning. To comprehend an argument and fill the gap between claims and reasons, it is vital to find the implicit supporting warrants behind. …
-
Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks
2018
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the …
-
Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network
2019 · arXiv (Cornell University)
Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we …
-
The Microsoft 2016 Conversational Speech Recognition System
2016 · arXiv (Cornell University)
We describe the 2017 version of Microsoft's conversational speech recognition system, in which we update our 2016 system with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art …
-
Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding …
-
Towards Faithful Neural Table-to-Text Generation with Content-Matching Constraints
2020
Text generation from a knowledge base aims to translate knowledge triples to naturallanguage descriptions. Most existing methods ignore the faithfulness between a generated text description and the original table, leading to generated information that goes …
-
Towards Robust Speaker Verification with Target Speaker Enhancement
2021 · arXiv (Cornell University)
This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV). Specifically, an enrollment speaker conditioned …
-
Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization …
-
Towards Improved Zero-shot Voice Conversion with Conditional DSVAE
2022 · arXiv (Cornell University)
Disentangling content and speaking style information is essential for zero-shot non-parallel voice conversion (VC). Our previous study investigated a novel framework with disentangled sequential variational autoencoder (DSVAE) as the backbone for information decomposition. We have …
-
Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse
2023
Self-supervised learning (SSL) models confront challenges of abrupt informational collapse or slow dimensional collapse. We propose TriNet, which introduces a novel triple-branch architecture for preventing collapse and stabilizing the pretraining. TriNet learns the SSL latent …
-
TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs
2023 · arXiv (Cornell University)
Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess …
-
Preference Alignment Improves Language Model-Based TTS
2025
Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences …
-
DocBench: A Benchmark for Evaluating LLM-based Document Reading Systems
2025
Anni Zou, Wenhao Yu, Hongming Zhang, Kaixin Ma, Deng Cai, Zhuosheng Zhang, Hai Zhao, Dong Yu. Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing. 2025.
-
DREAM: A Challenge Data Set and Models for Dialogue-Based Reading Comprehension
2019 · Transactions of the Association for Computational Linguistics
We present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data …
-
Investigating End-to-end Speech Recognition for Mandarin-english Code-switching
2019
Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, …