Hirofumi Inaguma
6 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
ESPnet-ST: All-in-One Speech Translation Toolkit
2020 · arXiv (Cornell University)
We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly implements automatic …
-
End-to-end speech-to-dialog-act recognition
2020 · arXiv (Cornell University)
Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by …
-
Enhancing Speech-To-Speech Translation with Multiple TTS Targets
2023
It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
2023 · arXiv (Cornell University)
Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to leverage …
-
SSR: Alignment-Aware Modality Connector for Speech Language Models
2025
Fusing speech into a pre-trained language model (SpeechLM) usually suffers from the inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality.We propose SSR-CONNECTOR (Segmented Speech Representation Connector) for better modality fusion.Leveraging speech-text …
-
A Comparative Study on Transformer vs RNN in Speech Applications
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art …