Zhilin Yang
10 papers in the PaperMetrix corpus
Papers by this author
-
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
2018 · arXiv (Cornell University)
Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) …
-
P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks
2022
Prompt tuning, which only tunes continuous prompts with a frozen language model, substantially reduces per-task storage and memory usage at training. However, in the context of NLU, prior work reveals that prompt tuning does not …
-
GPS: Genetic Prompt Search for Efficient Few-Shot Learning
2022
Prompt-based techniques have demostrated great potential for improving the few-shot generalization of pretrained language models. However, their performance heavily relies on the manual design of prompts and thus requiring a lot of human efforts. In …
-
Multi-Task Cross-Lingual Sequence Tagging from Scratch
2016 · arXiv (Cornell University)
We present a deep hierarchical recurrent neural network for sequence tagging. Given a sequence of words, our model employs deep gated recurrent units on both character and word levels to encode morphology and context information, …
-
Transfer Learning for Sequence Tagging with Hierarchical Recurrent Networks
2017 · arXiv (Cornell University)
Recent papers have shown that neural networks obtain state-of-the-art performance on several different sequence tagging tasks. One appealing property of such systems is their generality, as excellent performance can be achieved with a unified architecture …
-
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
2017 · arXiv (Cornell University)
We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is …
-
XLNet: Generalized Autoregressive Pretraining for Language Understanding
2019 · arXiv (Cornell University)
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency …
-
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
2019
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed …
-
P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
2021 · arXiv (Cornell University)
Prompt tuning, which only tunes continuous prompts with a frozen language model, substantially reduces per-task storage and memory usage at training. However, in the context of NLU, prior work reveals that prompt tuning does not …
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, Jie Tang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.