Varieties of Chinese Discrimination Using Hybrid Squeeze-and-Excitation Network
At a glance
- Citations
- 0
- References
- 24
- Comments
- 0
Abstract
As linguistic differences among related languages are less obvious than those among different languages, automatic language identification in similar languages, varieties, and dialects is a much more challenging task. In this paper, we propose a hybrid SENet (Squeeze-and-Excitation network) based attention model to discriminate varieties of Chinese from six regions (i.e., Mandarin Chinese, Hong Kong, Taiwan, Macao, Malaysia, and Singapore). Firstly, we encode the word embeddings using Bi-LSTM, and automatically extract the discriminative words related to different varieties of Chinese based on SENet. Then, since the original word-embeddings can implicitly encode rich linguistic regularities and patterns, a max-pooling strategy is adopted to extract the most salient features. Finally, we successfully incorporate the SENet and max-pooling to conduct varieties of Chinese discrimination. Experiments on our 6-way varieties of Chinese corpus show the effectiveness of the proposed model, which significantly outperforms the state-of-the-art widely used methods by a large margin.
Publication details
- DOI
- 10.1145/3377713.3377797
- OpenAlex
- W3004802554
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.