Inx-Speakerhub: A 2000-Hour Indian Multiligual Speaker Verification Corpus
At a glance
- الاستشهادات
- 0
- المراجع
- 13
- Comments
- 0
Abstract
This paper presents the data collection efforts, statistics and preparation involved in creating the 2333-hour INX-SpeakerHub, an Indian multilingual speaker identification dataset. It has legally collected speech from approximately 11,000 Indian native speakers in 10 different Indian languages. Until now, the VoxCeleb dataset (1+2) has been the most popular corpus used in many current state-of-theart systems. Therefore, VoxCeleb based speaker embedding extractors are often used by default even for Indian language-based speech applications. However, the proportion of Indian language data in VoxCeleb is quite less and might lead to subpar performance in speech tasks involving Indian languages. India is a country with 22 official languages and is home to 1.43 billion people. So creating a dataset like INX-SpeakerHub to build speaker embedding extractors for the Indian languages is of great interest. As our analysis shows, the VoxCeleb dataset has higher % equal error rate (% EER) in comparison to the INX-SpeakerHub for the speaker verification task in Indian languages. Further, we also analyse the speaker verification performance across different language families as well as on unseen languages.
Publication details
- DOI
- 10.1109/slt61566.2024.10832221
- OpenAlex
- W4406461691
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.