Researcher profile

Tatsuya Kawahara

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model training

    2016

    This paper addresses unsupervised training of DNN acoustic model, by exploiting a large amount of unlabeled data with CRF-based classifiers. In the proposed scheme, we obtain ASR hypotheses by complementary GMM and DNN based ASR …

  2. Automatic meeting transcription system for the Japanese parliament (diet)

    2017

    Applications of automatic speech recognition (ASR) have been extended to a variety of tasks and domains, including spontaneous human-human speech. We have developed an ASR system for the Japanese Pariliament (Diet), which has been operated …

  3. End-to-end speech-to-dialog-act recognition

    2020 · arXiv (Cornell University)

    Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by …

  4. A multi-party attentive listening robot which stimulates involvement from side participants

    2021

    We demonstrate the moderating abilities of a multi-party attentive listening robot system when multiple people are speaking in turns. Our conventional one-on-one attentive listening system generates listener responses such as backchannels, repeats, elaborating questions, and …

  5. Synthesizing waveform sequence-to-sequence to augment training data for sequence-to-sequence speech recognition

    2021 · Nippon Onkyo Gakkaishi/Acoustical science and technology/Nihon Onkyo Gakkaishi

    Sequence-to-sequence (seq2seq) automatic speech recognition (ASR) recently achieves state-of-the-art performance with fast decoding and a simple architecture. On the other hand, it requires a large amount of training data and cannot use text-only data for …

  6. Refining Synthesized Speech Using Speaker Information and Phone Masking for Data Augmentation of Speech Recognition

    2024 · IEEE/ACM Transactions on Audio Speech and Language Processing

    While end-to-end automatic speech recognition (ASR) has shown impressive performance, it requires a huge amount of speech and transcription data. The conversion of domain-matched text to speech (TTS) has been investigated as one approach to …

  7. StyEmp: Stylizing Empathetic Response Generation via Multi-Grained Prefix Encoder and Personality Reinforcement

    2024 · arXiv (Cornell University)

    Recent approaches for empathetic response generation mainly focus on emotional resonance and user understanding, without considering the system's personality. Consistent personality is evident in real human expression and is important for creating trustworthy systems. To …