Zhehuai Chen
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
An Asynchronous WFST-Based Decoder For Automatic Speech Recognition
2021 · arXiv (Cornell University)
We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which …
-
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024 · arXiv (Cornell University)
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is …
-
Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
2025
Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech …
-
EMMeTT: Efficient Multimodal Machine Translation Training
2025
A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal …