conference-paper

Voice spoofing detection with raw waveform based on Dual Path Res2net

Research footprint

At a glance

Citations
2
References
25
Comments
0
Paper overview

Abstract

The natural-sounding speech produced by recent text-to-speech and voice conversion techniques pose serious threats to automatic speaker verification systems. The majority of existing spoofing detection countermeasures perform well when the nature of the attacks is known during training. However, their performance in realistic applications degrades in dealing with unseen types of attacks. To address this concern, we propose a novel method for spoof detection, namely Dual Path Res2Net (DP-Res2Net) to improve the robustness to unknown attacks. As to the feature engineering, we employ the time domain features rather than the commonly-used frequency domain ones. We directly input the time domain features of 80,000 sampling points into the network. The input features are further processed by shallow feature learning module, interactive feature learning module, deep feature learning module as well as the discriminator network. The dual-path residual-like block exploit the dependence between successive pieces of audios with large receptive fields. Furthermore, the proposed DP-Res2Net significantly improves the model’s generalizability to unseen spoofing attacks. We evaluate the performance of the proposed method over public-available ASVspoof 2019 logic access evaluation set, and the results demonstrate that it outperforms state-of-the-art audio spoof detection models.

Record transparency

Publication details

DOI
10.1145/3503181.3503218
OpenAlex
W4214909824
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.