Researcher profile

Ralf Schlüter

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. The RWTH/UPB system combination for the CHiME 2018 Workshop

    2018

    This paper describes the systems for the single-array track and the multiple-array track of the 5th CHiME Challenge. The final system is a combination of multiple systems, using Confusion Network Combination (CNC). The different systems …

  2. Investigations on Phoneme-Based End-To-End Speech Recognition.

    2020 · arXiv (Cornell University)

    Common end-to-end models like CTC or encoder-decoder-attention models use characters or subword units like BPE as the output labels. We do systematic comparisons between grapheme-based and phoneme-based output labels. These can be single phonemes without …

  3. On Language Model Integration for RNN Transducer based Speech Recognition

    2021 · arXiv (Cornell University)

    The mismatch between an external language model (LM) and the implicitly learned internal LM (ILM) of RNN-Transducer (RNN-T) can limit the performance of LM integration such as simple shallow fusion. A Bayesian interpretation suggests to …

  4. Efficient Utilization of Large Pre-Trained Models for Low Resource ASR

    2023

    Unsupervised representation learning has recently helped automatic speech recognition (ASR) to tackle tasks with limited labeled data. Following this, hardware limitations and applications give rise to the question how to take advantage of large pre-trained …

  5. From Feedforward to Recurrent LSTM Neural Networks for Language Modeling

    2015 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Language models have traditionally been estimated based on relative frequencies, using count statistics that can be extracted from huge amounts of text data. More recently, it has been found that neural networks are particularly powerful …

  6. Improved Training of End-to-end Attention Models for Speech Recognition

    2018

    Sequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition. In this work, we show that such models can achieve competitive results on the Switchboard 300h and LibriSpeech 1000h tasks. In particular, we …