preprint Open access

HyperToken: Script-Aware, Nukta-Preserving Tokenization for Low-Resource Indic Languages

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Standard subword tokenizers are typically evaluated on aggregate compression metrics, which can obscure script-specific failure modes that disproportionately affect low-resource languages. We present HyperToken, a script-aware tokenizer that combines abugida syllable clustering with unigram vocabulary pruning, and compare it against PU-Tok (PolyglotUnigramTokenizer), a general-purpose polyglot baseline, on 30 typologically diverse languages from FLORES-200. Both tokenizers are trained under identical conditions (same corpus, same 24,000-token vocabulary budget, same held-out evaluation split). HyperToken achieves lower tokens/character than PU-Tok on 26 of 30 languages, including all nine Indic abugidas evaluated, and achieves perfect round-trip reconstruction (decode(encode(text)) == text) on all 30 languages. PU-Tok, by contrast, fails round-trip on 8 of 30 languages, concentrated in scripts that use nukta diacritics to represent loanword phonology. We argue that round-trip correctness should be reported as a first-class metric alongside compression when evaluating multilingual tokenizers.This deposit contains the paper, LaTeX source, the HyperToken tokenizer package, the PU-Tok baseline implementation, benchmark scripts, and raw/parsed results needed to reproduce all reported numbers.

Record transparency

Publication details

DOI
10.5281/zenodo.21624238
OpenAlex
W7171417942
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.