article Open access

Unsupervised Learning of Cross-Lingual Symbol Embeddings Without Parallel Data

  • ScholarWorks@UMassAmherst (University of Massachusetts Amherst)
  • University of Massachusetts Amherst
Research footprint

At a glance

Citations
1
References
29
Comments
0
Paper overview

Öz

We present a new method for unsupervised learning of multilingual symbol (e.g. character) embeddings, without any parallel data or prior knowledge about correspondences between languages. It is able to exploit similarities across languages between the distributions over symbols' contexts of use within their language, even in the absence of any symbols in common to the two languages. In experiments with an artificially corrupted text corpus, we show that the method can retrieve character correspondences obscured by noise. We then present encouraging results of applying the method to real linguistic data, including for low-resourced languages. The learned representations open the possibility of fully unsupervised comparative studies of text or speech corpora in low-resourced languages with no prior knowledge regarding their symbol sets.

Record transparency

Publication details

DOI
10.7275/wx64-ea83
OpenAlex
W2903381383
Document type
article
Language
EN
Source
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.