Zhengtao Yu
8 papers in the PaperMetrix corpus
Papers by this author
-
Lao Named Entity Recognition based on conditional random fields with simple heuristic information
2015
According to characteristics of Lao named entities, the paper proposes an approach of Lao Named Entity Recognition (NER) based on Conditional Random Fields (CRFs) with knowledge information. Firstly, we segment the text into word sequence …
-
Syntax-Based Chinese-Vietnamese Tree-to-Tree Statistical Machine Translation with Bilingual Features
2019 · ACM Transactions on Asian and Low-Resource Language Information Processing
Because of the scarcity of bilingual corpora, current Chinese--Vietnamese machine translation is far from satisfactory. Considering the differences between Chinese and Vietnamese, we investigate whether linguistic differences can be used to supervise machine translation and …
-
A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS Tagging
2021 · ACM Transactions on Asian and Low-Resource Language Information Processing
The smallest semantic unit of the Burmese language is called the syllable. In the present study, it is intended to propose the first neural joint learning model for Burmese syllable segmentation, word segmentation, and part-of-speech …
-
Feature-opinion pair identification method in two-stage based on dependency constraints
2018 · International Journal of Information and Communication Technology
Feature-opinion pair identification includes opinion words, opinion targets extraction and their relations identification, is important for analysis online reviews. In this paper, we propose a feature-opinion pair identification method in two-stage based on dependency constraints …
-
PTEKC: Pre-training with Event knowledge of ConceptNet for Cross-lingual Event Causality Identification
2024 · Research Square
Abstract Event Causality Identification (ECI) task aims to identify causal relations between events in texts, which is beneficial to understanding the precise logic meaning expressed by a text. Although existing event causality identification works based …
-
StreamingDialogue: Prolonged Dialogue Learning via Long Context Compression with Minimal Losses
2024 · arXiv (Cornell University)
Standard Large Language Models (LLMs) struggle with handling dialogues with long contexts due to efficiency and consistency issues. According to our observation, dialogue contexts are highly structured, and the special token of \textit{End-of-Utterance} (EoU) in …
-
Zero-Shot Text Normalization via Cross-Lingual Knowledge Distillation
2024 · IEEE/ACM Transactions on Audio Speech and Language Processing
Text normalization (TN) is a crucial preprocessing step in text-to-speech synthesis, which pertains to the accurate pronunciation of numbers and symbols within the text. Existing neural network-based TN methods have shown significant success in rich-resource …
-
2D-TPE: Two-Dimensional Positional Encoding Enhances Table Understanding for Large Language Models
2024 · arXiv (Cornell University)
Tables are ubiquitous across various domains for concisely representing structured information. Empowering large language models (LLMs) to reason over tabular data represents an actively explored direction. However, since typical LLMs only support one-dimensional~(1D) inputs, existing …