preprint وصول مفتوح

Feature-Rich Named Entity Recognition for Bulgarian Using Conditional Random Fields

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
22
المراجع
25
Comments
0
Paper overview

Abstract

The paper presents a feature-rich approach to the automatic recognition and categorization of named entities (persons, organizations, locations, and miscellaneous) in news text for Bulgarian. We combine well-established features used for other languages with language-specific lexical, syntactic and morphological information. In particular, we make use of the rich tagset annotation of the BulTreeBank (680 morpho-syntactic tags), from which we derive suitable task-specific tagsets (local and nonlocal). We further add domain-specific gazetteers and additional unlabeled data, achieving F1=89.4%, which is comparable to the state-of-the-art results for English.

Record transparency

Publication details

DOI
10.48550/arxiv.2109.15121
OpenAlex
W2251163391
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.