preprint Open access

Recognizing Multi-talker Speech with Permutation Invariant Training

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
12
References
31
Comments
0
Paper overview

Abstract

In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) for automatic speech recognition (ASR). In PIT-ASR, we compute the average cross entropy (CE) over all frames in the whole utterance for each possible output-target assignment, pick the one with the minimum CE, and optimize for that assignment. PIT-ASR forces all the frames of the same speaker to be aligned with the same output layer. This strategy elegantly solves the label permutation problem and speaker tracing problem in one shot. Our experiments on artificially mixed AMI data showed that the proposed approach is very promising.

Record transparency

Publication details

DOI
10.48550/arxiv.1704.01985
OpenAlex
W2606136275
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.