Constantine Lignos
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling
2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task. We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings-words …
-
What changes when you randomly choose BPE merge operations? Not much.
2023
We introduce two simple randomized variants of byte pair encoding (BPE) and explore whether randomizing the selection of merge operations substantially affects a downstream machine translation task. We focus on translation into morphologically rich languages, …
-
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
2024 · arXiv (Cornell University)
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize …
-
MasakhaNER: Named Entity Recognition for African Languages
2021 · Transactions of the Association for Computational Linguistics
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …