conference-paper

An Extensive Review: Models For Regional Language Speech Recognition

Research footprint

At a glance

الاستشهادات
0
المراجع
16
Comments
0
Paper overview

Abstract

Speech is the most common form of communication. Speech recognition remains an important milestone as it facilitates human-computer interaction and makes the communication rather simpler. It remains an asset across endless domains including healthcare, education, robotics and automation. Automatic Speech Recognition (ASR) systems have advanced significantly in their performance over the years, but transcribing languages remains a challenge. Hindi as a language has more than 600M speakers worldwide. The involvement of ASR systems is still far from the reach of these people. The primary reason for this challenge is the scattered and unorganised nature of available data. Our research aims to provide an analysis of mainstream and recognised speech recognition models for Hindi language: Wav2Vec2, Whisper, and DeepSpeech. Furthermore, our paper highlights each model’s strengths and limitations by diving into their architecture, learning methodologies, and unique characteristics. Noteworthy aspects of each model include Wav2Vec2’s self-supervised learning, Whisper’s reliance on multilingual datasets, and Deep Speech’s scalability in industry applications. By carefully weighing the pros and cons of each model, we intend to provide an analysis of the current status of advancements towards automatic speech recognition.

Record transparency

Publication details

DOI
10.1109/asiancon62057.2024.10837903
OpenAlex
W4406658816
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.