Post-Processing Tokenizer Adaptation Under a Fixed Vocabulary Budget: Compression–Performance Trade-offs Across Models and Tasks
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Subword tokenization is a key component of modern language models, but general-purpose tokenizers are often inefficient for domain-specific text . Recent work has proposed post-processing replacement methods that adapt a tokenizer by replacing low-frequency subwords with high-frequency words under a fixed vocabulary constraint. However, it remains unclear whether this strategy generalizes across different model architectures and tasks. In this paper, we study a fixed-vocabulary post-processing replacement strategy on two model architectures, GPT-2 and RoBERTa, and three text classification datasets: IMDB, SST-2, and AG News. We evaluate multiple replacement budgets and examine the trade-off between token reduction and classification accuracy. Our results show that the method is highly setting-dependent. GPT-2 maintains stable accuracy under moderate compression, while RoBERTa is more sensitive to replacement and generally shows lower accuracy than the baseline. Across datasets, SST-2 is the most sensitive to token changes, whereas AG News is relatively robust but gains little compression. These findings suggest that lightweight tokenizer adaptation is possible, but its effectiveness depends strongly on model architecture and task characteristics, and compression alone is not sufficient to predict downstream behavior.
Publication details
- DOI
- 10.47297/wspciwsp2516-252743.20261002
- OpenAlex
- W7167051110
- Document type
- article
- Language
- EN
- Source
- Creativity and Innovation
- Last metadata update
Comments
Log in to join the discussion.