Zihang Dai
7 papers in the PaperMetrix corpus
Papers by this author
-
CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge Bases
2016
How can we enable computers to automatically answer questions like "Who created the character Harry Potter"? Carefully built knowledge bases provide rich sources of facts. However, it remains a challenge to answer factoid questions raised …
-
A Mutual Information Maximization Perspective of Language Representation Learning
2020 · arXiv (Cornell University)
We show state-of-the-art word representation learning methods maximize an objective function that is a lower bound on the mutual information between different parts of a word sequence (i.e., a sentence). Our formulation provides an alternative …
-
An Interpretable Knowledge Transfer Model for Knowledge Base Completion
2017 · arXiv (Cornell University)
Knowledge bases are important resources for a variety of natural language processing tasks but suffer from incompleteness. We propose a novel embedding model, \emph{ITransF}, to perform knowledge base completion. Equipped with a sparse attention mechanism, …
-
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
2017 · arXiv (Cornell University)
We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is …
-
SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation
2018
In this work, we examine methods for data augmentation for text-based tasks such as neural machine translation (NMT). We formulate the design of a data augmentation policy with desirable properties as an optimization problem, and …
-
XLNet: Generalized Autoregressive Pretraining for Language Understanding
2019 · arXiv (Cornell University)
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency …
-
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
2019
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed …