HyperToken: Script-Aware, Nukta-Preserving Tokenization for Low-Resource Indic Languages
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Standard subword tokenizers are typically evaluated on aggregate compression metrics, which can obscure script-specific failure modes that disproportionately affect low-resource languages. We present HyperToken, a script-aware tokenizer that combines abugida syllable clustering with unigram vocabulary pruning, and compare it against PU-Tok (PolyglotUnigramTokenizer), a general-purpose polyglot baseline, on 30 typologically diverse languages from FLORES-200. Both tokenizers are trained under identical conditions (same corpus, same 24,000-token vocabulary budget, same held-out evaluation split). HyperToken achieves lower tokens/character than PU-Tok on 26 of 30 languages, including all nine Indic abugidas evaluated, and achieves perfect round-trip reconstruction (decode(encode(text)) == text) on all 30 languages. PU-Tok, by contrast, fails round-trip on 8 of 30 languages, concentrated in scripts that use nukta diacritics to represent loanword phonology. We argue that round-trip correctness should be reported as a first-class metric alongside compression when evaluating multilingual tokenizers.This deposit contains the paper, LaTeX source, the HyperToken tokenizer package, the PU-Tok baseline implementation, benchmark scripts, and raw/parsed results needed to reproduce all reported numbers.
Publication details
- DOI
- 10.5281/zenodo.21624238
- OpenAlex
- W7171417942
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.