Researcher profile

Jason Phang

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks

    2018 · arXiv (Cornell University)

    Pretraining sentence encoders with language modeling and related unsupervised tasks has recently been shown to be very effective for language understanding tasks. By supplementing language model-style pretraining with further training on data-rich supervised tasks, such …

  2. Do Attention Heads in BERT Track Syntactic Dependencies?

    2019 · arXiv (Cornell University)

    We investigate the extent to which individual attention heads in pretrained transformer language models, such as BERT and RoBERTa, implicitly capture syntactic dependency relations. We employ two methods---taking the maximum attention weight and computing the …

  3. Intermediate-Task Transfer Learning with Pretrained Models for Natural Language Understanding: When and Why Does It Work?

    2020 · arXiv (Cornell University)

    While pretrained models such as BERT have shown large gains across natural language understanding tasks, their performance can be improved by further training the model on a data-rich intermediate task, before fine-tuning it on a …

  4. Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?

    2020

    Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  5. The Pile: An 800GB Dataset of Diverse Text for Language Modeling

    2020 · arXiv (Cornell University)

    Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus …