Damai Dai
4 papers in the PaperMetrix corpus
Papers by this author
-
Neural Knowledge Bank for Pretrained Transformers
2022 · arXiv (Cornell University)
The ability of pretrained Transformers to remember factual knowledge is essential but still limited for existing models. Inspired by existing work that regards Feed-Forward Networks (FFNs) in Transformers as key-value memories, we design a Neural …
-
Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and Optimization
2023 · arXiv (Cornell University)
Pretrained language models have achieved remarkable success in natural language understanding. However, fine-tuning pretrained models on limited training data tends to overfit and thus diminish performance. This paper presents Bi-Drop, a fine-tuning strategy that selectively …
-
Exploring Activation Patterns of Parameters in Language Models
2024 · arXiv (Cornell University)
Most work treats large language models as black boxes without in-depth understanding of their internal working mechanism. In order to explain the internal representations of LLMs, we propose a gradient-based metric to assess the activation …
-
A Survey on In-context Learning
2022 · arXiv (Cornell University)
With the increasing capabilities of large language models (LLMs), in-context learning (ICL) has emerged as a new paradigm for natural language processing (NLP), where LLMs make predictions based on contexts augmented with a few examples. …