Orevaoghene Ahia
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
2024 · arXiv (Cornell University)
Identifying linguistic differences between dialects of a language often requires expert knowledge and meticulous human analysis. This is largely due to the complexity and nuance involved in studying various dialects. We present a novel approach …
-
FLEXITOKENS: Flexible Tokenization for Evolving Language Models
2026
Adapting language models to new data distributions by simple finetuning is challenging.This is due to the rigidity of their subword tokenizers, which typically remain unchanged during adaptation.This inflexibility often leads to inefficient tokenization, causing overfragmentation …
-
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
2022 · Transactions of the Association for Computational Linguistics
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality …
-
MasakhaNER: Named Entity Recognition for African Languages
2021 · Transactions of the Association for Computational Linguistics
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …