Researcher profile

Hirofumi Inaguma

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. ESPnet-ST: All-in-One Speech Translation Toolkit

    2020 · arXiv (Cornell University)

    We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly implements automatic …

  2. End-to-end speech-to-dialog-act recognition

    2020 · arXiv (Cornell University)

    Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by …

  3. Enhancing Speech-To-Speech Translation with Multiple TTS Targets

    2023

    It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …

  4. Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

    2023 · arXiv (Cornell University)

    Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to leverage …

  5. SSR: Alignment-Aware Modality Connector for Speech Language Models

    2025

    Fusing speech into a pre-trained language model (SpeechLM) usually suffers from the inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality.We propose SSR-CONNECTOR (Segmented Speech Representation Connector) for better modality fusion.Leveraging speech-text …

  6. A Comparative Study on Transformer vs RNN in Speech Applications

    2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …