Modeling a Hybrid Stemmer for Kokborok
At a glance
- Citations
- 3
- References
- 15
- Comments
- 0
Öz
Kokborok is a resource scarce and vulnerable language and is spoken only by around a million people in the north-east Indian state of Tripura. Lots of unstructured textual data in Kokborok is now available however efficient data mining for such a resource is not applicable. Reasons are the absence of an efficient pre-processing step that can help in solving the information retrieval or natural language processing. In this work, we present a model of hybrid stemming algorithm for the vulnerable and resource-scarce Kokborok language and assess its performance using various methods. Evaluation of the hybrid stemmer was performed on a corpus of 52,482 words that resulted in achieving an accuracy of 89.28%. Recall and Precision of 89.28% and 89.61 %, respectively, with F1-measure of 89.44 has been found for the stemmer.
Publication details
- DOI
- 10.1109/isacc56298.2023.10084267
- OpenAlex
- W4362496889
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.