Weizhu Chen
16 papers in the PaperMetrix corpus
Papers by this author
-
ReasoNet: Learning to Stop Reading in Machine Comprehension
2016 · arXiv (Cornell University)
Teaching a computer to read and answer general questions pertaining to a document is a challenging yet unsolved problem. In this paper, we describe a novel neural network architecture called the Reasoning Network (ReasoNet) for …
-
ARCH: Efficient Adversarial Regularized Training with Caching
2021
Adversarial regularization can improve model generalization in many natural language processing tasks. However, conventional approaches are computationally expensive since they need to generate a perturbation for each sample in each epoch. We propose a new …
-
Token-wise Curriculum Learning for Neural Machine Translation
2021 · arXiv (Cornell University)
Existing curriculum learning approaches to Neural Machine Translation (NMT) require sampling sufficient amounts of "easy" samples from training data at the early training stage. This is not always achievable for low-resource languages where the amount …
-
Controllable Natural Language Generation with Contrastive Prefixes
2022 · Findings of the Association for Computational Linguistics: ACL 2022
To guide the generation of large pretrained language models (LM), previous work has focused on directly fine-tuning the language model or utilizing an attribute discriminator. In this work, we propose a novel lightweight framework for …
-
CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight Sharing
2022 · arXiv (Cornell University)
Model ensemble is a popular approach to produce a low-variance and well-generalized model. However, it induces large memory and inference costs, which are often not affordable for real-world deployment. Existing work has resorted to sharing …
-
Meet in the Middle: A New Pre-training Paradigm
2023 · arXiv (Cornell University)
Most language models (LMs) are trained and applied in an autoregressive left-to-right fashion, assuming that the next token only depends on the preceding ones. However, this assumption ignores the potential benefits of using the full …
-
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
2023 · arXiv (Cornell University)
Large language models have made significant progress in various language tasks, yet they still struggle with complex mathematics. In this paper, we propose ToRA a series of Tool-integrated Reasoning Agents designed to solve challenging mathematical …
-
MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning
2024 · arXiv (Cornell University)
Parameter-efficient fine-tuning (PEFT) has been widely employed for domain adaptation, with LoRA being one of the most prominent methods due to its simplicity and effectiveness. However, in multi-task learning (MTL) scenarios, LoRA tends to obscure …
-
Multi-Task Deep Neural Networks for Natural Language Understanding
2019 · arXiv (Cornell University)
In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a …
-
Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
2019 · arXiv (Cornell University)
This paper explores the use of knowledge distillation to improve a Multi-Task Deep Neural Network (MT-DNN) (Liu et al., 2019) for learning text representations across multiple natural language understanding tasks. Although ensemble learning can improve …
-
Adversarial Training for Large Neural Language Models
2020 · arXiv (Cornell University)
Generalization and robustness are both key desiderata for designing machine learning methods. Adversarial training can enhance robustness, but past work often finds it hurts generalization. In natural language processing (NLP), pre-training large neural language models …
-
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
2020 · arXiv (Cornell University)
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture DeBERTa (Decoding-enhanced BERT with disentangled attention) that …
-
What Makes Good In-Context Examples for GPT-3?
2022
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Chen. Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures. 2022.
-
DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION
2021 · International Conference on Learning Representations
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture \textbf{DeBERTa} (\textbf{D}ecoding-\textbf{e}nhanced \textbf{BERT} with disentangled \textbf{a}ttention) that …
-
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
2021 · arXiv (Cornell University)
This paper presents a new pre-trained language model, DeBERTaV3, which improves the original DeBERTa model by replacing mask language modeling (MLM) with replaced token detection (RTD), a more sample-efficient pre-training task. Our analysis shows that …
-
Making Large Language Models Better Reasoners with Step-Aware Verifier
2022 · arXiv (Cornell University)
Few-shot learning is a challenging task that requires language models to generalize from limited examples. Large language models like GPT-3 and PaLM have made impressive progress in this area, but they still face difficulties in …