conference-paper

Study of the Architectural Features of the ResNet Neural Network Model for Solving the Task of Speaker Recognition

Research footprint

At a glance

الاستشهادات
0
المراجع
12
Comments
0
Paper overview

Abstract

The relevance of the use of speech recognition methods is due to the growing need for security, convenience, personalisation, efficiency and automation of business processes in various aspects of modern life. The authors can also mention the most modern factor that further increases the relevance of the study of speaker recognition methods protection against falsification and transcription of voice messages in social and communication platforms. The relevance of the identification and verification problem determines the variety of methods used to solve the problem from characteristics determined by the physical nature of sound to neural network models of deep learning. Considering the problem area and the subject of the study, the authors managed to offer a detailed systematization of speaker identification and verification approaches. As a result of the study, a generalized speaker recognition model based on the ResNet neural network model with convolutional architecture was proposed, for experimental testing of which a training dataset based on speeches of five speakers was prepared and restructured. Analysis of the model components revealed the possibility of improvement by accelerating the Fast Fourier Transform on a massively parallel computing system. This improvement showed the possibility of speeding up computations by up to 1,469 times even for small audio files (22.881 MB), using the data parallelism available in modern discrete GPUs, as well as the features of the FFT algorithm. The analysis of the relevance and demand for the problem of identification and verification has led to the creation of a detailed systematisation of approaches to speaker identification and verification, which clearly visualises the methods of signal processing, their purpose and the nature of the data. The study of the architectural features of the ResNet neural network model with convolutional architecture for speaker recognition showed an accuracy of 0.9747, achieved after 14 epochs of model training. After the 14th epoch, the model stops learning, as evidenced by the increase in the value of validation losses. The further development of the work is an in-depth analysis of speaker identification, namely, improving the model by the criterion of resistance to various speech characteristics, such as different accents, intonations, speech styles, etc. For this purpose, it is planned to revise the training set and expand it with different noise levels and different types of environment to assess the model's resistance to external factors.

Record transparency

Publication details

DOI
10.1109/khpiweek61434.2024.10877951
OpenAlex
W4407638195
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.