Frequency and temporal convolutional attention for text-independent\n speaker recognition
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Öz
Majority of the recent approaches for text-independent speaker recognition\napply attention or similar techniques for aggregation of frame-level feature\ndescriptors generated by a deep neural network (DNN) front-end. In this paper,\nwe propose methods of convolutional attention for independently modelling\ntemporal and frequency information in a convolutional neural network (CNN)\nbased front-end. Our system utilizes convolutional block attention modules\n(CBAMs) [1] appropriately modified to accommodate spectrogram inputs. The\nproposed CNN front-end fitted with the proposed convolutional attention modules\noutperform the no-attention and spatial-CBAM baselines by a significant margin\non the VoxCeleb [2, 3] speaker verification benchmark, and our best model\nachieves an equal error rate of 2:031% on the VoxCeleb1 test set, improving the\nexisting state of the art result by a significant margin. For a more thorough\nassessment of the effects of frequency and temporal attention in real-world\nconditions, we conduct ablation experiments by randomly dropping frequency bins\nand temporal frames from the input spectrograms, concluding that instead of\nmodelling either of the entities, simultaneously modelling temporal and\nfrequency attention translates to better real-world performance.\n
Publication details
- DOI
- 10.48550/arxiv.1910.07364
- OpenAlex
- W4288091711
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Oturum Açın to join the discussion.