preprint وصول مفتوح

Speaker Clustering With Neural Networks And Audio Processing

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
4
المراجع
1
Comments
0
Paper overview

Abstract

Speaker clustering is the task of differentiating speakers in a recording. In a way, the aim is to answer "who spoke when" in audio recordings. A common method used in industry is feature extraction directly from the recording thanks to MFCC features, and by using well-known techniques such as Gaussian Mixture Models (GMM) and Hidden Markov Models (HMM). In this paper, we studied neural networks (especially CNN) followed by clustering and audio processing in the quest to reach similar accuracy to state-of-the-art methods.

Record transparency

Publication details

DOI
10.48550/arxiv.1803.08276
OpenAlex
W2794073214
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.