conference-paper وصول مفتوح

UDON: Unsupervised Data SelectiON for Biomedical Entity Recognition

Research footprint

At a glance

الاستشهادات
1
المراجع
23
Comments
0
Paper overview

Abstract

High-quality training datasets are critical for building successful Machine Learning (ML) based NLP systems. However, these datasets are not always available in low-resource contexts such as the biomedical domain. Here, selecting relevant training data is as important as the choice of the ML model. In this study we propose UDON: Unsupervised Data selectiON for biomedical entity recognition using domain-specific pretrained Language Models (LMs). We first show that pretrained LMs succeed at implicitly learning the differences between datasets without any supervision, and then use these models to select relevant data instances. Next, we evaluate the proposed methods for entity recognition on seven biomedical datasets and one news domain dataset using four LMs and three selection methods. Our results show that using pretrained domain-specific LMs for data selection outperforms all other approaches. Finally, we use domain classification as an auxiliary task for pretraining the neural network on the in-domain dataset and show this yields further improvements.

Record transparency

Publication details

DOI
10.1145/3507524.3507525
OpenAlex
W4225838254
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.