conference-paper
Open access
Data representation methods and use of mined corpora for Indian language transliteration
Research footprint
At a glance
- Citations
- 15
- References
- 16
- Comments
- 0
Paper overview
Öz
Our NEWS 2015 shared task submission is a PBSMT based transliteration system with the following corpus preprocessing enhancements: (i) addition of wordboundary markers, and (ii) languageindependent, overlapping character segmentation. We show that the addition of word-boundary markers improves transliteration accuracy substantially, whereas our overlapping segmentation shows promise in our preliminary analysis. We also compare transliteration systems trained using manually created corpora with the ones mined from parallel translation corpus for English to Indian language pairs. We identify the major errors in English to Indian language transliterations by analyzing heat maps of confusion matrices.
Record transparency
Publication details
- DOI
- 10.18653/v1/w15-3912
- OpenAlex
- W2250943466
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.