ملف الباحث
Yingfa Chen
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
SHUOWEN-JIEZI: Linguistically Informed Tokenizers For Chinese Language Model Pretraining
2021 · arXiv (Cornell University)
Conventional tokenization methods for Chinese pretrained language models (PLMs) treat each character as an indivisible token (Devlin et al., 2019), which ignores the characteristics of the Chinese writing system. In this work, we comprehensively study …