Researcher profile

Jiatong Shi

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization …

  2. EURO: ESPnet Unsupervised ASR Open-source Toolkit

    2022 · arXiv (Cornell University)

    This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, …

  3. Enhancing Speech-To-Speech Translation with Multiple TTS Targets

    2023

    It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …

  4. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

    2024 · arXiv (Cornell University)

    Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is …

  5. Preference Alignment Improves Language Model-Based TTS

    2025

    Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences …

  6. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

    2025 · arXiv (Cornell University)

    Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses …

  7. SUPERB: Speech Processing Universal PERformance Benchmark

    2021

    Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various …