preprint
Open access
Massively Multilingual Word Embeddings
Research footprint
At a glance
- Citations
- 282
- References
- 33
- Comments
- 0
Paper overview
Abstract
We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do not require parallel data. Our new evaluation method, multiQVEC-CCA, is shown to correlate better than previous ones with two downstream tasks (text categorization and parsing). We also describe a web portal for evaluation that will facilitate further research in this area, along with open-source releases of all our methods.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1602.01925
- OpenAlex
- W2270364989
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.