Hung-yi Lee
19 papers in the PaperMetrix corpus
Papers by this author
-
Towards End-to-end Speech-to-text Translation with Two-pass Decoding
2019
Speech-to-text translation (ST) refers to transforming the audio in source language to the text in target language. Mainstream solutions for such tasks are to cascade automatic speech recognition with machine translation, for which the transcriptions …
-
Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal Classification
2017
Connectionist temporal classification (CTC) is a powerful approach for sequence-to-sequence learning, and has been popularly used in speech recognition.The central ideas of CTC include adding a label "blank" during training.With this mechanism, CTC eliminates the …
-
Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
2020
In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by …
-
Auto-KWS 2021 Challenge: Task, Datasets, and Baselines
2021
Auto-KWS 2021 challenge calls for automated machine learning (AutoML) solutions to automate the process of applying machine learning to a customized keyword spotting task.Compared with other keyword spotting tasks, Auto-KWS challenge has the following three …
-
Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models
2021 · arXiv (Cornell University)
Code-switching (CS) is common in daily conversations where more than one language is used within a sentence. The difficulties of CS speech recognition lie in alternating languages and the lack of transcribed data. Therefore, this …
-
Membership Inference Attacks Against Self-supervised Speech Models
2022 · Interspeech 2022
Recently, adapting the idea of self-supervised learning (SSL) on continuous speech has started gaining attention.SSL models pre-trained on a huge amount of unlabeled audio can generate general-purpose representations that benefit a wide variety of speech …
-
Unsupervised Multiple Choices Question Answering: Start Learning from Basic Knowledge
2020 · arXiv (Cornell University)
In this paper, we study the possibility of almost unsupervised Multiple Choices Question Answering (MCQA). Starting from very basic knowledge, MCQA model knows that some choices have higher probabilities of being correct than the others. …
-
Few Shot Cross-Lingual TTS Using Transferable Phoneme Embedding
2022 · Interspeech 2022
This paper studies a transferable phoneme embedding framework that aims to deal with the cross-lingual text-to-speech (TTS) problem under the few-shot setting.Transfer learning is a common approach when it comes to few-shot learning since training …
-
Exploring Efficient-tuning Methods in Self-supervised Speech Models
2022 · arXiv (Cornell University)
In this study, we aim to explore efficient tuning methods for speech self-supervised learning. Recent studies show that self-supervised learning (SSL) can learn powerful representations for different speech tasks. However, fine-tuning pre-trained models for each …
-
EURO: ESPnet Unsupervised ASR Open-source Toolkit
2022 · arXiv (Cornell University)
This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, …
-
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
2024 · arXiv (Cornell University)
Large Language Models (LLMs) have demonstrated effectiveness not only in language tasks but also in video reasoning. This paper introduces a novel dataset, Tropes in Movies (TiM), designed as a testbed for exploring two critical …
-
Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
2024 · arXiv (Cornell University)
The diffusion model, a new generative modeling paradigm, has achieved significant success in generating images, audio, video, and text. It has been adapted for sequence-to-sequence text generation (Seq2Seq) through DiffuSeq, termed S2S Diffusion. Existing S2S-Diffusion …
-
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024 · arXiv (Cornell University)
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is …
-
Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
2025
Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech …
-
Supervised and Unsupervised Transfer Learning for Question Answering
2018
Yu-An Chung, Hung-Yi Lee, James Glass. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
-
DyKgChat: Benchmarking Dialogue Generation Grounding on Dynamic Knowledge Graphs
2019
Yi-Lin Tuan, Yun-Nung Chen, Hung-yi Lee. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Code-Switching Sentence Generation by Generative Adversarial Networks and its Application to Data Augmentation
2019
Code-switching is about dealing with alternative languages in speech or text.It is partially speaker-dependent and domainrelated, so completely explaining the phenomenon by linguistic rules is challenging.Compared to most monolingual tasks, insufficient data is an issue …
-
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
2021 · IEEE/ACM Transactions on Audio Speech and Language Processing
We introduce a self-supervised speech pre-training method called TERA, which stands for Transformer Encoder Representations from Alteration. Recent approaches often learn by using a single auxiliary task like contrastive prediction, autoregressive prediction, or masked reconstruction. …
-
SUPERB: Speech Processing Universal PERformance Benchmark
2021
Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various …