conference-paper Open access

Joint Embeddings of Chinese Words, Characters, and Fine-grained Subcharacter Components

Research footprint

At a glance

Citations
122
References
24
Comments
0
Paper overview

Abstract

Word embeddings have attracted much attention recently. Different from alphabetic writing systems, Chinese characters are often composed of subcharacter components which are also semantically informative. In this work, we propose an approach to jointly embed Chinese words as well as their characters and fine-grained subcharacter components. We use three likelihoods to evaluate whether the context words, characters, and components can predict the current target word, and collected 13,253 subcharacter components to demonstrate the existing approaches of decomposing Chinese characters are not enough. Evaluation on both word similarity and word analogy tasks demonstrates the superior performance of our model.

Record transparency

Publication details

DOI
10.18653/v1/d17-1027
OpenAlex
W2759366113
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.