Zhifang Sui
9 papers in the PaperMetrix corpus
Papers by this author
-
Table-to-Text Generation by Structure-Aware Seq2seq Learning
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Table-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a …
-
Incorporating Glosses into Neural Word Sense Disambiguation
2018
Word Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context. Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods. …
-
Neural Knowledge Bank for Pretrained Transformers
2022 · arXiv (Cornell University)
The ability of pretrained Transformers to remember factual knowledge is essential but still limited for existing models. Inspired by existing work that regards Feed-Forward Networks (FFNs) in Transformers as key-value memories, we design a Neural …
-
RepCL: Exploring Effective Representation for Continual Text Classification
2023 · arXiv (Cornell University)
Continual learning (CL) aims to constantly learn new knowledge over time while avoiding catastrophic forgetting on old tasks. In this work, we focus on continual text classification under the class-incremental setting. Recent CL studies find …
-
Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and Optimization
2023 · arXiv (Cornell University)
Pretrained language models have achieved remarkable success in natural language understanding. However, fine-tuning pretrained models on limited training data tends to overfit and thus diminish performance. This paper presents Bi-Drop, a fine-tuning strategy that selectively …
-
Exploring Activation Patterns of Parameters in Language Models
2024 · arXiv (Cornell University)
Most work treats large language models as black boxes without in-depth understanding of their internal working mechanism. In order to explain the internal representations of LLMs, we propose a gradient-based metric to assess the activation …
-
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
2025 · arXiv (Cornell University)
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, …
-
Jointly Extracting Event Triggers and Arguments by Dependency-Bridge RNN and Tensor-Based Argument Interaction
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Event extraction plays an important role in natural language processing (NLP) applications including question answering and information retrieval. Traditional event extraction relies heavily on lexical and syntactic features, which require intensive human engineering and may …
-
A Survey on In-context Learning
2022 · arXiv (Cornell University)
With the increasing capabilities of large language models (LLMs), in-context learning (ICL) has emerged as a new paradigm for natural language processing (NLP), where LLMs make predictions based on contexts augmented with a few examples. …