ملف الباحث

Zihang Dai

7 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge Bases

    2016

    How can we enable computers to automatically answer questions like "Who created the character Harry Potter"? Carefully built knowledge bases provide rich sources of facts. However, it remains a challenge to answer factoid questions raised …

  2. A Mutual Information Maximization Perspective of Language Representation Learning

    2020 · arXiv (Cornell University)

    We show state-of-the-art word representation learning methods maximize an objective function that is a lower bound on the mutual information between different parts of a word sequence (i.e., a sentence). Our formulation provides an alternative …

  3. An Interpretable Knowledge Transfer Model for Knowledge Base Completion

    2017 · arXiv (Cornell University)

    Knowledge bases are important resources for a variety of natural language processing tasks but suffer from incompleteness. We propose a novel embedding model, \emph{ITransF}, to perform knowledge base completion. Equipped with a sparse attention mechanism, …

  4. Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

    2017 · arXiv (Cornell University)

    We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is …

  5. SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation

    2018

    In this work, we examine methods for data augmentation for text-based tasks such as neural machine translation (NMT). We formulate the design of a data augmentation policy with desirable properties as an optimization problem, and …

  6. XLNet: Generalized Autoregressive Pretraining for Language Understanding

    2019 · arXiv (Cornell University)

    With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency …

  7. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

    2019

    Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed …