Researcher profile

Shinji Watanabe

26 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling

    2018 · arXiv (Cornell University)

    Sequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech research. The approach benefits by performing model training without using lexicon and alignments. However, this poses a new problem of requiring more …

  2. Attention-based ASR with Lightweight and Dynamic Convolutions

    2019 · arXiv (Cornell University)

    End-to-end (E2E) automatic speech recognition (ASR) with sequence-to-sequence models has gained attention because of its simple model training compared with conventional hidden Markov model based ASR. Recently, several studies report the state-of-the-art E2E ASR results …

  3. ESPnet-ST: All-in-One Speech Translation Toolkit

    2020 · arXiv (Cornell University)

    We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly implements automatic …

  4. Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition

    2021

    Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the …

  5. End-To-End Speaker Diarization as Post-Processing

    2021

    This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech …

  6. Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models

    2021 · arXiv (Cornell University)

    Non-autoregressive (NAR) modeling has gained more and more attention in speech processing. With recent state-of-the-art attention-based automatic speech recognition (ASR) structure, NAR can realize promising real-time factor (RTF) improvement with only small degradation of accuracy …

  7. On Prosody Modeling for ASR+TTS Based Voice Conversion

    2021 · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to transcribe the source speech into the underlying …

  8. Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors

    2022 · arXiv (Cornell University)

    A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label …

  9. End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation

    2022 · arXiv (Cornell University)

    This work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech Input for Self-supervised learning representation (IRIS). Compared with conventional E2E ASR models, …

  10. BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model

    2022 · arXiv (Cornell University)

    This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC). Our formulation relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through …

  11. EURO: ESPnet Unsupervised ASR Open-source Toolkit

    2022 · arXiv (Cornell University)

    This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, …

  12. Enhancing Speech-To-Speech Translation with Multiple TTS Targets

    2023

    It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …

  13. Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam Search

    2024

    End-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired …

  14. Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss

    2024 · arXiv (Cornell University)

    Contextualized end-to-end automatic speech recognition has been an active research area, with recent efforts focusing on the implicit learning of contextual phrases based on the final loss objective. However, these approaches ignore the useful contextual …

  15. SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data

    2024 · arXiv (Cornell University)

    In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous research that focused on lip motion as visual …

  16. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

    2024 · arXiv (Cornell University)

    Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is …

  17. Preference Alignment Improves Language Model-Based TTS

    2025

    Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences …

  18. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    2025 · arXiv (Cornell University)

    The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a large-scale …

  19. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

    2025 · arXiv (Cornell University)

    Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, existing E2E approaches often require large-scale training data and generates responses …

  20. DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

    2025 · arXiv (Cornell University)

    We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing distillation methods that rely on layer-wise and feature-wise mapping between teacher and student models, …

  21. Joint CTC-attention based end-to-end speech recognition using multi-task learning

    2017

    Recently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping between variable-length input …

  22. ESPnet: End-to-End Speech Processing Toolkit

    2018 · arXiv (Cornell University)

    This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a …

  23. Cycle-consistency Training for End-to-end Speech Recognition

    2019

    This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, …

  24. A Comparative Study on Transformer vs RNN in Speech Applications

    2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …

  25. Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

    2020

    This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and …

  26. SUPERB: Speech Processing Universal PERformance Benchmark

    2021

    Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various …