Kyunghyun Cho
42 papers in the PaperMetrix corpus
Papers by this author
-
Query-Efficient Imitation Learning for End-to-End Autonomous Driving
2016 · arXiv (Cornell University)
One way to approach end-to-end autonomous driving is to learn a policy function that maps from a sensory input, such as an image frame from a front-facing camera, to a driving action, by imitating an …
-
Larger-Context Language Modelling with Recurrent Neural Network
2016
In this work, we propose a novel method to incorporate corpus-level discourse information into language modelling. We call this larger-context language model. We introduce a late fusion approach to a recurrent language model based on …
-
Training a Ranking Function for Open-Domain Question Answering
2018 · arXiv (Cornell University)
In recent years, there have been amazing advances in deep learning methods for machine reading. In machine reading, the machine reader has to extract the answer from the given ground truth paragraph. Recently, the state-of-the-art …
-
BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
2019 · arXiv (Cornell University)
We show that BERT (Devlin et al., 2018) is a Markov random field language model. This formulation gives way to a natural procedure to sample sentences from BERT. We generate from BERT and find that …
-
Learning to Understand Phrases by Embedding the Dictionary
2015 · arXiv (Cornell University)
Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …
-
Dialogue Natural Language Inference
2019
Consistency is a long standing issue faced by dialogue models. In this paper, we frame the consistency of dialogue agents as natural language inference (NLI) and create a new natural language inference dataset called Dialogue …
-
Evaluation of Combined Artificial Intelligence and Radiologist Assessment to Interpret Screening Mammograms
2020 · JAMA Network Open
Importance: Mammography screening currently relies on subjective human interpretation. Artificial intelligence (AI) advances could be used to increase mammography screening accuracy by reducing missed cancers and false positives. Objective: To evaluate whether AI can overcome …
-
The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event Prediction
2021 · arXiv (Cornell University)
Event schemas encode knowledge of stereotypical structures of events and their connections. As events unfold, schemas are crucial to act as a scaffolding. Previous work on event schema induction focuses either on atomic events or …
-
xVal: A Continuous Numerical Tokenization for Scientific Language Models
2023 · arXiv (Cornell University)
Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate …
-
Language Models as Causal Effect Generators
2024 · arXiv (Cornell University)
In this work, we present sequence-driven structural causal models (SD-SCMs), a framework for specifying causal models with user-defined structure and language-model-defined mechanisms. We characterize how an SD-SCM enables sampling from observational, interventional, and counterfactual distributions …
-
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
2025 · arXiv (Cornell University)
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented -- enabling smaller student models to emulate …
-
Learning to Understand Phrases by Embedding the Dictionary
2016 · Transactions of the Association for Computational Linguistics
Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …
-
Gated Feedback Recurrent Neural Networks
2015 · arXiv (Cornell University)
In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper …
-
On Using Monolingual Corpora in Neural Machine Translation
2015 · HAL (Le Centre pour la Communication Scientifique Directe)
Recent work on end-to-end neural network-based architectures for machine translation has shown promising results for En-Fr and En-De translation. Arguably, one of the major factors behind this success has been the availability of high quality …
-
A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition
2015
Deep neural networks have advanced the state-of-the-art in automatic speech recognition, when combined with hidden Markov models (HMMs). Recently there has been interest in using systems based on recurrent neural networks (RNNs) to perform sequence …
-
On Using Very Large Target Vocabulary for Neural Machine Translation
2015
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism
2016
We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. …
-
Montreal Neural Machine Translation Systems for WMT’15
2015
Neural machine translation (NMT) systems have recently achieved results comparable to the state of the art on a few translation tasks, including EnglishFrench and EnglishGerman. The main purpose of the Montreal Institute for Learning Algorithms …
-
Efficient Character-level Document Classification by Combining Convolution and Recurrent Layers
2016 · arXiv (Cornell University)
Document classification tasks were primarily tackled at word level. Recent research that works with character-level inputs shows several benefits over word-level approaches such as natural incorporation of morphemes and better handling of rare words. We …
-
Learning Distributed Representations of Sentences from Unlabelled Data
2016
Unsupervised methods for learning distributed representations of words are ubiquitous in today's NLP research, but far less is known about the best ways to learn distributed phrase or sentence representations from unlabelled data. This paper …
-
A Character-level Decoder without Explicit Segmentation for Neural Machine Translation
2016
The existing machine translation systems, whether phrase-based or neural, have relied almost exclusively on word-level modelling with explicit segmentation. In this paper, we ask a fundamental question: can neural machine translation generate a character sequence …
-
Zero-Resource Translation with Multi-Lingual Neural Machine Translation
2016 · arXiv (Cornell University)
In this paper, we propose a novel finetuning algorithm for the recently introduced multi-way, mulitlingual neural machine translate that enables zero-resource machine translation. When used together with novel many-to-one translation strategies, we empirically show that …
-
Fully Character-Level Neural Machine Translation without Explicit Segmentation
2017 · Transactions of the Association for Computational Linguistics
Most existing machine translation systems operate at the level of words, relying on explicit segmentation to extract tokens. We introduce a neural machine translation (NMT) model that maps a source character sequence to a target …
-
Learning to Parse and Translate Improves Neural Machine Translation
2017
There has been relatively little attention to incorporating linguistic prior to neural machine translation. Much of the previous work was further constrained to considering linguistic prior on the source side. In this paper, we propose …
-
Nematus: a Toolkit for Neural Machine Translation
2017
Rico Sennrich, Orhan Firat, Kyunghyun Cho, Alexandra Birch, Barry Haddow, Julian Hitschler, Marcin Junczys-Dowmunt, Samuel Läubli, Antonio Valerio Miceli Barone, Jozef Mokry, Maria Nădejde. Proceedings of the Software Demonstrations of the 15th Conference of the …
-
SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
2017 · arXiv (Cornell University)
We publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect a full pipeline of …
-
Search Engine Guided Neural Machine Translation
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
In this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. …
-
Code-Switched Named Entity Recognition with Embedding Attention
2018
We describe our work for the CALCS 2018 shared task on named entity recognition on code-switched data. Our system ranked first place for MS Arabic-Egyptian named entity recognition and third place for English-Spanish.
-
Meta-Learning for Low-Resource Neural Machine Translation
2018
In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT). We frame low-resource translation as a metalearning problem, and we learn …
-
Passage Re-ranking with BERT
2019 · arXiv (Cornell University)
Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural …
-
Non-Monotonic Sequential Text Generation
2019 · arXiv (Cornell University)
Standard sequential generation methods assume a pre-specified generation order, such as text generation methods which generate words from left to right. In this work, we propose a framework for training models of text generation that …
-
Document Expansion by Query Prediction
2019 · arXiv (Cornell University)
One technique to improve the retrieval effectiveness of a search engine is to expand documents with terms that are related or representative of the documents' content.From the perspective of a question answering system, this might …
-
Attention-Based Models for Speech Recognition
2015 · arXiv (Cornell University)
Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend …
-
Zero-Shot Transfer Learning for Event Extraction
2018
Most previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort. We take a fresh look at event extraction …
-
Task-Oriented Query Reformulation with Reinforcement Learning
2017
Search engines play an important role in our everyday lives by assisting us in finding the information we need. When we input a complex query, however, results are often far from satisfactory. In this work, …
-
Dynamic Meta-Embeddings for Improved Sentence Representations
2018
While one of the first steps in many NLP systems is selecting what pre-trained word embeddings to use, we argue that such a step is better left for neural networks to figure out by themselves. …
-
Neural Text Generation with Unlikelihood Training
2019 · arXiv (Cornell University)
Neural text generation is a key tool in natural language applications, but it is well known there are major problems at its core. In particular, standard likelihood training and decoding leads to dull and repetitive …
-
Multi-Stage Document Ranking with BERT
2019 · arXiv (Cornell University)
The advent of deep neural networks pre-trained via language modeling tasks has spurred a number of successful applications in natural language processing. This work explores one such popular model, BERT, in the context of document …
-
Insertion-based Decoding with Automatically Inferred Generation Order
2019 · Transactions of the Association for Computational Linguistics
Conventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. …
-
Neural Machine Translation with Byte-Level Subwords
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
Almost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up …
-
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
2020
Practical applications of abstractive summarization models are limited by frequent factual inconsistencies with respect to their input. Existing automatic evaluation metrics for summarization are largely insensitive to such errors. We propose QAGS, 1 an automatic …
-
True Few-Shot Learning with Language Models
2021 · arXiv (Cornell University)
Pretrained language models (LMs) perform well on many tasks even when learning from a few examples, but prior work uses many held-out examples to tune various aspects of learning, such as hyperparameters, training objectives, and …