Researcher profile
Prashant ARYA
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
HyperToken: Script-Aware, Nukta-Preserving Tokenization for Low-Resource Indic Languages
2026 · Zenodo (CERN European Organization for Nuclear Research)
Standard subword tokenizers are typically evaluated on aggregate compression metrics, which can obscure script-specific failure modes that disproportionately affect low-resource languages. We present HyperToken, a script-aware tokenizer that combines abugida syllable clustering with unigram vocabulary …