conference-paper وصول مفتوح

A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations

  • Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Research footprint

At a glance

الاستشهادات
14
المراجع
40
Comments
0
Paper overview

Abstract

Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models. We explore this problem from a novel angle of geometric algebra and semantic space. A simple but highly effective method "Language Information Removal (LIR)" factors out language identity information from semantic related components in multilingual representations pre-trained on multi-monolingual data. A post-training and model-agnostic method, LIR only uses simple linear operations, e.g. matrix factorization and orthogonal projection. LIR reveals that for weak-alignment multilingual systems, the principal components of semantic spaces primarily encodes language identity information. We first evaluate the LIR on a cross-lingual question answer retrieval task (LAReQA), which requires the strong alignment for the multilingual embedding space. Experiment shows that LIR is highly effectively on this task, yielding almost 100% relative improvement in MAP for weakalignment models. We then evaluate the LIR on Amazon Reviews and XEVAL dataset, with the observation that removing language information is able to improve the cross-lingual transfer performance.

Record transparency

Publication details

DOI
10.18653/v1/2021.emnlp-main.470
OpenAlex
W3200726972
Document type
conference-paper
Language
EN
Source
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.