preprint وصول مفتوح

Updating Pre-trained Word Vectors and Text Classifiers using Monolingual Alignment

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
8
المراجع
11
Comments
0
Paper overview

Abstract

In this paper, we focus on the problem of adapting word vector-based models to new textual data. Given a model pre-trained on large reference data, how can we adapt it to a smaller piece of data with a slightly different language distribution? We frame the adaptation problem as a monolingual word vector alignment problem, and simply average models after alignment. We align vectors using the RCSLS criterion. Our formulation results in a simple and efficient algorithm that allows adapting general-purpose models to changing word distributions. In our evaluation, we consider applications to word embedding and text classification models. We show that the proposed approach yields good performance in all setups and outperforms a baseline consisting in fine-tuning the model on new data.

Record transparency

Publication details

DOI
10.48550/arxiv.1910.06241
OpenAlex
W2980022889
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.