conference-paper

Inx-Speakerhub: A 2000-Hour Indian Multiligual Speaker Verification Corpus

Research footprint

At a glance

Citations
0
References
13
Comments
0
Paper overview

Abstract

This paper presents the data collection efforts, statistics and preparation involved in creating the 2333-hour INX-SpeakerHub, an Indian multilingual speaker identification dataset. It has legally collected speech from approximately 11,000 Indian native speakers in 10 different Indian languages. Until now, the VoxCeleb dataset (1+2) has been the most popular corpus used in many current state-of-theart systems. Therefore, VoxCeleb based speaker embedding extractors are often used by default even for Indian language-based speech applications. However, the proportion of Indian language data in VoxCeleb is quite less and might lead to subpar performance in speech tasks involving Indian languages. India is a country with 22 official languages and is home to 1.43 billion people. So creating a dataset like INX-SpeakerHub to build speaker embedding extractors for the Indian languages is of great interest. As our analysis shows, the VoxCeleb dataset has higher % equal error rate (% EER) in comparison to the INX-SpeakerHub for the speaker verification task in Indian languages. Further, we also analyse the speaker verification performance across different language families as well as on unseen languages.

Record transparency

Publication details

DOI
10.1109/slt61566.2024.10832221
OpenAlex
W4406461691
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.