Researcher profile

Yun Tang

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Multilingual Speech Translation with Efficient Finetuning of Pretrained Models

    2020 · arXiv (Cornell University)

    We present a simple yet effective approach to build multilingual speech-to-text (ST) translation by efficient transfer learning from pretrained speech encoder and text decoder. Our key finding is that a minimalistic LNA (LayerNorm and Attention) …

  2. A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks

    2021

    Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily relies on the availability of large amounts of training data. This …

  3. Enhancing Speech-To-Speech Translation with Multiple TTS Targets

    2023

    It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …

  4. Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

    2023 · arXiv (Cornell University)

    Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to leverage …

  5. Transducer Consistency Regularization For Speech to Text Applications

    2024

    Consistency regularization is a commonly used practice to encourage the model to generate consistent representation from distorted input features and improve model generalization. It shows significant improvement on various speech applications that are optimized with …

  6. Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization

    2025 · arXiv (Cornell University)

    Low latency speech human-machine communication is becoming increasingly necessary as speech technology advances quickly in the last decade. One of the primary factors behind the advancement of speech technology is self-supervised learning. Most self-supervised learning …

  7. Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs

    2019

    Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the …