ملف الباحث

Shengxiang Gao

5 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Lao Named Entity Recognition based on conditional random fields with simple heuristic information

    2015

    According to characteristics of Lao named entities, the paper proposes an approach of Lao Named Entity Recognition (NER) based on Conditional Random Fields (CRFs) with knowledge information. Firstly, we segment the text into word sequence …

  2. Syntax-Based Chinese-Vietnamese Tree-to-Tree Statistical Machine Translation with Bilingual Features

    2019 · ACM Transactions on Asian and Low-Resource Language Information Processing

    Because of the scarcity of bilingual corpora, current Chinese--Vietnamese machine translation is far from satisfactory. Considering the differences between Chinese and Vietnamese, we investigate whether linguistic differences can be used to supervise machine translation and …

  3. A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS Tagging

    2021 · ACM Transactions on Asian and Low-Resource Language Information Processing

    The smallest semantic unit of the Burmese language is called the syllable. In the present study, it is intended to propose the first neural joint learning model for Burmese syllable segmentation, word segmentation, and part-of-speech …

  4. PTEKC: Pre-training with Event knowledge of ConceptNet for Cross-lingual Event Causality Identification

    2024 · Research Square

    Abstract Event Causality Identification (ECI) task aims to identify causal relations between events in texts, which is beneficial to understanding the precise logic meaning expressed by a text. Although existing event causality identification works based …

  5. Zero-Shot Text Normalization via Cross-Lingual Knowledge Distillation

    2024 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Text normalization (TN) is a crucial preprocessing step in text-to-speech synthesis, which pertains to the accurate pronunciation of numbers and symbols within the text. Existing neural network-based TN methods have shown significant success in rich-resource …