preprint Open access

Comparative Analysis of Audio Features for Unsupervised Speaker Change Detection

  • Preprints.org
Research footprint

At a glance

Citations
2
References
0
Comments
0
Paper overview

Abstract

This study analyzes how various audio features influence speaker change detection in different unsupervised approaches. Ten different audio features, including MFCC, mel-spectrogram, spectral centroid, and pitch, etc. were chosen for analysis. Two distinct unsupervised methods were selected for comparing the effectiveness of these feature: the Bayesian information criterion with Gaussian mixture model (BIC-GMM), a model-based approach, and the Kullback-Leibler divergence with Gaussian (KL-GMM), a metric-based approach. In order to gain a prior insight into which features have a significant impact on the SCD problem, we calculate statistic for two probabilities: the probability of speaker change given there is a feature change, and vice versa. To experimentally check the prior insights, a set of experiments carried out to analyze the influence of various audio feature for SCD. The results showed that in both methods, MFCC demonstrates a robust performance, and ranks at the top. Zero crossing rate, chroma and spectral contrast gives a higher performance in BIC-GMM. Mel-spectrogram consistently ranks at the bottom, suggesting it has the least impact on performance across both methods.

Record transparency

Publication details

DOI
10.20944/preprints202410.1209.v1
OpenAlex
W4403454546
Document type
preprint
Language
EN
Source
Preprints.org
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.