conference-paper

A Chinese-Thai Cross-language Word Embedding Method Based on Unequal Corpus of Small Dictionaries

Research footprint

At a glance

الاستشهادات
0
المراجع
8
Comments
0
Paper overview

Abstract

This paper proposes a Chinese-Thailand cross-language word embedding method based on the unequal corpus of a small dictionary. The method first normalizes the word vectors of Chinese and the small dictionary, and obtains the gradient descent for the orthogonal optimal linear transformation of the small dictionary words. The initial value is then clustered on a large Chinese corpus, with the help of a small dictionary to find the Chinese word vector corresponding to each cluster cluster, and take the mean value of each cluster word vector obtained by the clustering and the mean value of the word vector corresponding to Chinese and Thai, Establish a new bilingual word vector correspondence, and extend the newly established bilingual word vector to the small dictionary, so that the Chinese-Thai small dictionary can be generalized and expanded. Finally, use the Chinese-Thai dictionary after generalization to perform gradient descent on the cross-language word embedding mapping model to obtain the optimal value.

Record transparency

Publication details

DOI
10.1109/bdee52938.2021.00035
OpenAlex
W4206274156
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.