Research on Russian Cultural Transliteration Algorithm Based on Hidden Markov Model
At a glance
- Citations
- 0
- References
- 16
- Comments
- 0
Abstract
Transliteration of cultural terminology poses unique challenges in natural language processing, particularly for the Russian language with its rich cultural lexicon. This research addresses these challenges by developing a transliteration algorithm tailored to Russian cultural terms using a Hidden Markov Model (HMM). The proposed approach enhances the Speech Assessment Methods Phonetic Alphabet (SAMPA)-based phoneme set to account for the phonetic intricacies of cultural vocabulary, where stress and vowel reduction patterns play crucial roles. A specialized Russian cultural pronunciation dictionary containing over 10,000 terms, encompassing historical, literary, and folk terminology is involved. The proposed algorithm employs a data-driven methodology, starting with a "many-to-many" alignment using the Expectation-Maximization algorithm to accommodate the diverse phonetic representations of cultural terms. The alignment is further refined through a joint N-gram model and encapsulated into a robust HMM, which effectively models the probabilistic nature of phonetic variations. The decoding phase of the HMM accurately predicts the pronunciation of complex cultural terms. The algorithm's efficacy was evaluated through meticulous cross-validation, yielding an average word-form accuracy of 71.2% and a phoneme accuracy of 89.7%, demonstrating its potential in enhancing the accessibility and preservation of Russian cultural heritage in digital formats.
Publication details
- DOI
- 10.1145/3662739.3664742
- OpenAlex
- W4401288557
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.