Researcher profile

Dong Yu

16 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Recognizing Multi-talker Speech with Permutation Invariant Training

    2017 · arXiv (Cornell University)

    In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) …

  2. BLCU_NLP at SemEval-2018 Task 12: An Ensemble Model for Argument Reasoning Based on Hierarchical Attention

    2018

    The argument comprehension reasoning task aims to reconstruct and analyze the argument reasoning. To comprehend an argument and fill the gap between claims and reasons, it is vital to find the implicit supporting warrants behind. …

  3. Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks

    2018

    Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the …

  4. Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network

    2019 · arXiv (Cornell University)

    Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we …

  5. The Microsoft 2016 Conversational Speech Recognition System

    2016 · arXiv (Cornell University)

    We describe the 2017 version of Microsoft's conversational speech recognition system, in which we update our 2016 system with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art …

  6. Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment

    2020 · Proceedings of the AAAI Conference on Artificial Intelligence

    Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding …

  7. Towards Faithful Neural Table-to-Text Generation with Content-Matching Constraints

    2020

    Text generation from a knowledge base aims to translate knowledge triples to naturallanguage descriptions. Most existing methods ignore the faithfulness between a generated text description and the original table, leading to generated information that goes …

  8. Towards Robust Speaker Verification with Target Speaker Enhancement

    2021 · arXiv (Cornell University)

    This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV). Specifically, an enrollment speaker conditioned …

  9. Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization …

  10. Towards Improved Zero-shot Voice Conversion with Conditional DSVAE

    2022 · arXiv (Cornell University)

    Disentangling content and speaking style information is essential for zero-shot non-parallel voice conversion (VC). Our previous study investigated a novel framework with disentangled sequential variational autoencoder (DSVAE) as the backbone for information decomposition. We have …

  11. Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse

    2023

    Self-supervised learning (SSL) models confront challenges of abrupt informational collapse or slow dimensional collapse. We propose TriNet, which introduces a novel triple-branch architecture for preventing collapse and stabilizing the pretraining. TriNet learns the SSL latent …

  12. TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs

    2023 · arXiv (Cornell University)

    Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess …

  13. Preference Alignment Improves Language Model-Based TTS

    2025

    Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences …

  14. DocBench: A Benchmark for Evaluating LLM-based Document Reading Systems

    2025

    Anni Zou, Wenhao Yu, Hongming Zhang, Kaixin Ma, Deng Cai, Zhuosheng Zhang, Hai Zhao, Dong Yu. Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing. 2025.

  15. DREAM: A Challenge Data Set and Models for Dialogue-Based Reading Comprehension

    2019 · Transactions of the Association for Computational Linguistics

    We present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data …

  16. Investigating End-to-end Speech Recognition for Mandarin-english Code-switching

    2019

    Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, …