conference-paper

Modeling a Hybrid Stemmer for Kokborok

Research footprint

At a glance

Citations
3
References
15
Comments
0
Paper overview

Abstract

Kokborok is a resource scarce and vulnerable language and is spoken only by around a million people in the north-east Indian state of Tripura. Lots of unstructured textual data in Kokborok is now available however efficient data mining for such a resource is not applicable. Reasons are the absence of an efficient pre-processing step that can help in solving the information retrieval or natural language processing. In this work, we present a model of hybrid stemming algorithm for the vulnerable and resource-scarce Kokborok language and assess its performance using various methods. Evaluation of the hybrid stemmer was performed on a corpus of 52,482 words that resulted in achieving an accuracy of 89.28%. Recall and Precision of 89.28% and 89.61 %, respectively, with F1-measure of 89.44 has been found for the stemmer.

Record transparency

Publication details

DOI
10.1109/isacc56298.2023.10084267
OpenAlex
W4362496889
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.