Researcher profile

Prashant ARYA

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. HyperToken: Script-Aware, Nukta-Preserving Tokenization for Low-Resource Indic Languages

    2026 · Zenodo (CERN European Organization for Nuclear Research)

    Standard subword tokenizers are typically evaluated on aggregate compression metrics, which can obscure script-specific failure modes that disproportionately affect low-resource languages. We present HyperToken, a script-aware tokenizer that combines abugida syllable clustering with unigram vocabulary …