conference-paper Open access

Subword-level Composition Functions for Learning Word Embeddings

Research footprint

At a glance

Citations
28
References
57
Comments
0
Paper overview

Öz

Subword-level information is crucial for capturing the meaning and morphology of words, especially for out-of-vocabulary entries. We propose CNN-and RNN-based subword-level composition functions for learning word embeddings, and systematically compare them with popular word-level and subword-level models (Skip-Gram and FastText). Additionally, we propose a hybrid training scheme in which a pure subword-level model is trained jointly with a conventional word-level embedding model based on lookup-tables. This increases the fitness of all types of subwordlevel word embeddings; the word-level embeddings can be discarded after training, leaving only compact subword-level representation with much smaller data volume. We evaluate these embeddings on a set of intrinsic and extrinsic tasks, showing that subwordlevel models have advantage on tasks related to morphology and datasets with high OOV rate, and can be combined with other types of embeddings.

Record transparency

Publication details

DOI
10.18653/v1/w18-1205
OpenAlex
W2807036468
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.