conference-paper

Research on Russian Cultural Transliteration Algorithm Based on Hidden Markov Model

Research footprint

At a glance

Citations
0
References
16
Comments
0
Paper overview

Abstract

Transliteration of cultural terminology poses unique challenges in natural language processing, particularly for the Russian language with its rich cultural lexicon. This research addresses these challenges by developing a transliteration algorithm tailored to Russian cultural terms using a Hidden Markov Model (HMM). The proposed approach enhances the Speech Assessment Methods Phonetic Alphabet (SAMPA)-based phoneme set to account for the phonetic intricacies of cultural vocabulary, where stress and vowel reduction patterns play crucial roles. A specialized Russian cultural pronunciation dictionary containing over 10,000 terms, encompassing historical, literary, and folk terminology is involved. The proposed algorithm employs a data-driven methodology, starting with a "many-to-many" alignment using the Expectation-Maximization algorithm to accommodate the diverse phonetic representations of cultural terms. The alignment is further refined through a joint N-gram model and encapsulated into a robust HMM, which effectively models the probabilistic nature of phonetic variations. The decoding phase of the HMM accurately predicts the pronunciation of complex cultural terms. The algorithm's efficacy was evaluated through meticulous cross-validation, yielding an average word-form accuracy of 71.2% and a phoneme accuracy of 89.7%, demonstrating its potential in enhancing the accessibility and preservation of Russian cultural heritage in digital formats.

Record transparency

Publication details

DOI
10.1145/3662739.3664742
OpenAlex
W4401288557
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.