article Open access

Post-Processing Tokenizer Adaptation Under a Fixed Vocabulary Budget: Compression–Performance Trade-offs Across Models and Tasks

  • Creativity and Innovation
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Subword tokenization is a key component of modern language models, but general-purpose tokenizers are often inefficient for domain-specific text . Recent work has proposed post-processing replacement methods that adapt a tokenizer by replacing low-frequency subwords with high-frequency words under a fixed vocabulary constraint. However, it remains unclear whether this strategy generalizes across different model architectures and tasks. In this paper, we study a fixed-vocabulary post-processing replacement strategy on two model architectures, GPT-2 and RoBERTa, and three text classification datasets: IMDB, SST-2, and AG News. We evaluate multiple replacement budgets and examine the trade-off between token reduction and classification accuracy. Our results show that the method is highly setting-dependent. GPT-2 maintains stable accuracy under moderate compression, while RoBERTa is more sensitive to replacement and generally shows lower accuracy than the baseline. Across datasets, SST-2 is the most sensitive to token changes, whereas AG News is relatively robust but gains little compression. These findings suggest that lightweight tokenizer adaptation is possible, but its effectiveness depends strongly on model architecture and task characteristics, and compression alone is not sufficient to predict downstream behavior.

Record transparency

Publication details

DOI
10.47297/wspciwsp2516-252743.20261002
OpenAlex
W7167051110
Document type
article
Language
EN
Source
Creativity and Innovation
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.