A Chinese-Thai Cross-language Word Embedding Method Based on Unequal Corpus of Small Dictionaries
At a glance
- Citations
- 0
- References
- 8
- Comments
- 0
Abstract
This paper proposes a Chinese-Thailand cross-language word embedding method based on the unequal corpus of a small dictionary. The method first normalizes the word vectors of Chinese and the small dictionary, and obtains the gradient descent for the orthogonal optimal linear transformation of the small dictionary words. The initial value is then clustered on a large Chinese corpus, with the help of a small dictionary to find the Chinese word vector corresponding to each cluster cluster, and take the mean value of each cluster word vector obtained by the clustering and the mean value of the word vector corresponding to Chinese and Thai, Establish a new bilingual word vector correspondence, and extend the newly established bilingual word vector to the small dictionary, so that the Chinese-Thai small dictionary can be generalized and expanded. Finally, use the Chinese-Thai dictionary after generalization to perform gradient descent on the cross-language word embedding mapping model to obtain the optimal value.
Publication details
- DOI
- 10.1109/bdee52938.2021.00035
- OpenAlex
- W4206274156
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.