conference-paper Open access

Token-level Identification of Multiword Expressions using Pre-trained Multilingual Language Models

Research footprint

At a glance

Citations
0
References
23
Comments
0
Paper overview

Abstract

In this paper, we consider novel cross-lingual settings for multiword expression (MWE) identification (Ramisch et al., 2020) and idiomaticity prediction (Tayyar Madabushi et al., 2022) in which systems are tested on languages that are unseen during training. Our findings indicate that pre-trained multilingual language models are able to learn knowledge about MWEs and idiomaticity that is not languagespecific. Moreover, we find that training data from other languages can be leveraged to give improvements over monolingual models.

Record transparency

Publication details

DOI
10.18653/v1/2023.mwe-1.1
OpenAlex
W4386576695
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.