Researcher profile

Heyan Huang

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Complicated Table Structure Recognition

    2019 · arXiv (Cornell University)

    The task of table structure recognition aims to recognize the internal structure of a table, which is a key step to make machines understand tables. Currently, there are lots of studies on this task for …

  2. Incorporating Domain Knowledge into Text Classification Diagnosis in Customer Service Dialogue Field

    2021 · Journal of Physics Conference Series

    Abstract The customer service dialogue process is an important way for consumers to communicate with manufacturers. In order to enhance the consumer experience as well as to assist the staff, we build a knowledge base …

  3. Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word Alignment

    2021 · arXiv (Cornell University)

    The cross-lingual language models are typically pretrained with masked language modeling on multilingual text or parallel sentences. In this paper, we introduce denoising word alignment as a new cross-lingual pre-training task. Specifically, the model first …

  4. Momentum Decoding: Open-ended Text Generation As Graph Exploration

    2022 · arXiv (Cornell University)

    Open-ended text generation with autoregressive language models (LMs) is one of the core tasks in natural language processing. However, maximization-based decoding methods (e.g., greedy/beam search) often lead to the degeneration problem, i.e., the generated text …

  5. SNER-CS: Self-training Named Entity Recognition in Computer Science

    2023 · Journal of Physics Conference Series

    Abstract As the number of scientific publications grows, especially in computer science domain (CS), it is important to extract scientific entities from a large number of CS publications. Distantly supervised methods, generating distantly annotated training …

  6. CSE: Conceptual Sentence Embeddings based on Attention Model

    2016

    Most sentence embedding models typically represent each sentence only using word surface, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance representation capability of sentence, we employ conceptualization model to …

  7. Jointly Multiple Events Extraction via Attention-based Graph Information Aggregation

    2018

    Event extraction is of practical utility in natural language processing. In the real world, it is a common phenomenon that multiple events existing in the same sentence, where extracting them are more difficult than extracting …