ملف الباحث

Zhifang Sui

9 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Table-to-Text Generation by Structure-Aware Seq2seq Learning

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Table-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a …

  2. Incorporating Glosses into Neural Word Sense Disambiguation

    2018

    Word Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context. Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods. …

  3. Neural Knowledge Bank for Pretrained Transformers

    2022 · arXiv (Cornell University)

    The ability of pretrained Transformers to remember factual knowledge is essential but still limited for existing models. Inspired by existing work that regards Feed-Forward Networks (FFNs) in Transformers as key-value memories, we design a Neural …

  4. RepCL: Exploring Effective Representation for Continual Text Classification

    2023 · arXiv (Cornell University)

    Continual learning (CL) aims to constantly learn new knowledge over time while avoiding catastrophic forgetting on old tasks. In this work, we focus on continual text classification under the class-incremental setting. Recent CL studies find …

  5. Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and Optimization

    2023 · arXiv (Cornell University)

    Pretrained language models have achieved remarkable success in natural language understanding. However, fine-tuning pretrained models on limited training data tends to overfit and thus diminish performance. This paper presents Bi-Drop, a fine-tuning strategy that selectively …

  6. Exploring Activation Patterns of Parameters in Language Models

    2024 · arXiv (Cornell University)

    Most work treats large language models as black boxes without in-depth understanding of their internal working mechanism. In order to explain the internal representations of LLMs, we propose a gradient-based metric to assess the activation …

  7. From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

    2025 · arXiv (Cornell University)

    Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, …

  8. Jointly Extracting Event Triggers and Arguments by Dependency-Bridge RNN and Tensor-Based Argument Interaction

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Event extraction plays an important role in natural language processing (NLP) applications including question answering and information retrieval. Traditional event extraction relies heavily on lexical and syntactic features, which require intensive human engineering and may …

  9. A Survey on In-context Learning

    2022 · arXiv (Cornell University)

    With the increasing capabilities of large language models (LLMs), in-context learning (ICL) has emerged as a new paradigm for natural language processing (NLP), where LLMs make predictions based on contexts augmented with a few examples. …