conference-paper

Relationship Between Speakers' Physiological Structure and Acoustic Speech Signals: Data-Driven Study Based on Frequency-Wise Attentional Neural Network

  • 2022 30th European Signal Processing Conference (EUSIPCO)
Research footprint

At a glance

Citations
2
References
29
Comments
0
Paper overview

Öz

Quantitatively revealing the relationship between speakers' physiological structure and acoustic speech signals by considering the properties of resonance and antiresonance can help us to extract effective speaker discriminative information (SDI) from speech signals. The conventional quantification method based on F-ratio only considers the power of acoustic speech in each frequency band independently. We propose a novel frequency-wise attentional neural network to learn the nonlinear combined effect of the frequency components on speaker identity. The learned results indicate that antiresonance frequency induced by the nasal cavity is another essential factor for speaker discrimination that the F-ratio method could not reveal. To further evaluate our findings, we designed a non-uniform subband processing strategy based on the learned results for speaker feature extraction and did automatic speaker verification (ASV). The ASV results confirmed that further emphasizing the spectral structure around the antiresonance frequency region can enhance speaker discrimination.

Record transparency

Publication details

DOI
10.23919/eusipco55093.2022.9909649
OpenAlex
W4312560113
Document type
conference-paper
Language
EN
Source
2022 30th European Signal Processing Conference (EUSIPCO)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.