ملف الباحث

Kyunghyun Cho

42 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Query-Efficient Imitation Learning for End-to-End Autonomous Driving

    2016 · arXiv (Cornell University)

    One way to approach end-to-end autonomous driving is to learn a policy function that maps from a sensory input, such as an image frame from a front-facing camera, to a driving action, by imitating an …

  2. Larger-Context Language Modelling with Recurrent Neural Network

    2016

    In this work, we propose a novel method to incorporate corpus-level discourse information into language modelling. We call this larger-context language model. We introduce a late fusion approach to a recurrent language model based on …

  3. Training a Ranking Function for Open-Domain Question Answering

    2018 · arXiv (Cornell University)

    In recent years, there have been amazing advances in deep learning methods for machine reading. In machine reading, the machine reader has to extract the answer from the given ground truth paragraph. Recently, the state-of-the-art …

  4. BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

    2019 · arXiv (Cornell University)

    We show that BERT (Devlin et al., 2018) is a Markov random field language model. This formulation gives way to a natural procedure to sample sentences from BERT. We generate from BERT and find that …

  5. Learning to Understand Phrases by Embedding the Dictionary

    2015 · arXiv (Cornell University)

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  6. Dialogue Natural Language Inference

    2019

    Consistency is a long standing issue faced by dialogue models. In this paper, we frame the consistency of dialogue agents as natural language inference (NLI) and create a new natural language inference dataset called Dialogue …

  7. Evaluation of Combined Artificial Intelligence and Radiologist Assessment to Interpret Screening Mammograms

    2020 · JAMA Network Open

    Importance: Mammography screening currently relies on subjective human interpretation. Artificial intelligence (AI) advances could be used to increase mammography screening accuracy by reducing missed cancers and false positives. Objective: To evaluate whether AI can overcome …

  8. The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event Prediction

    2021 · arXiv (Cornell University)

    Event schemas encode knowledge of stereotypical structures of events and their connections. As events unfold, schemas are crucial to act as a scaffolding. Previous work on event schema induction focuses either on atomic events or …

  9. xVal: A Continuous Numerical Tokenization for Scientific Language Models

    2023 · arXiv (Cornell University)

    Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate …

  10. Language Models as Causal Effect Generators

    2024 · arXiv (Cornell University)

    In this work, we present sequence-driven structural causal models (SD-SCMs), a framework for specifying causal models with user-defined structure and language-model-defined mechanisms. We characterize how an SD-SCM enables sampling from observational, interventional, and counterfactual distributions …

  11. Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation

    2025 · arXiv (Cornell University)

    Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented -- enabling smaller student models to emulate …

  12. Learning to Understand Phrases by Embedding the Dictionary

    2016 · Transactions of the Association for Computational Linguistics

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  13. Gated Feedback Recurrent Neural Networks

    2015 · arXiv (Cornell University)

    In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper …

  14. On Using Monolingual Corpora in Neural Machine Translation

    2015 · HAL (Le Centre pour la Communication Scientifique Directe)

    Recent work on end-to-end neural network-based architectures for machine translation has shown promising results for En-Fr and En-De translation. Arguably, one of the major factors behind this success has been the availability of high quality …

  15. A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition

    2015

    Deep neural networks have advanced the state-of-the-art in automatic speech recognition, when combined with hidden Markov models (HMMs). Recently there has been interest in using systems based on recurrent neural networks (RNNs) to perform sequence …

  16. On Using Very Large Target Vocabulary for Neural Machine Translation

    2015

    Sébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  17. Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism

    2016

    We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. …

  18. Montreal Neural Machine Translation Systems for WMT’15

    2015

    Neural machine translation (NMT) systems have recently achieved results comparable to the state of the art on a few translation tasks, including EnglishFrench and EnglishGerman. The main purpose of the Montreal Institute for Learning Algorithms …

  19. Efficient Character-level Document Classification by Combining Convolution and Recurrent Layers

    2016 · arXiv (Cornell University)

    Document classification tasks were primarily tackled at word level. Recent research that works with character-level inputs shows several benefits over word-level approaches such as natural incorporation of morphemes and better handling of rare words. We …

  20. Learning Distributed Representations of Sentences from Unlabelled Data

    2016

    Unsupervised methods for learning distributed representations of words are ubiquitous in today's NLP research, but far less is known about the best ways to learn distributed phrase or sentence representations from unlabelled data. This paper …

  21. A Character-level Decoder without Explicit Segmentation for Neural Machine Translation

    2016

    The existing machine translation systems, whether phrase-based or neural, have relied almost exclusively on word-level modelling with explicit segmentation. In this paper, we ask a fundamental question: can neural machine translation generate a character sequence …

  22. Zero-Resource Translation with Multi-Lingual Neural Machine Translation

    2016 · arXiv (Cornell University)

    In this paper, we propose a novel finetuning algorithm for the recently introduced multi-way, mulitlingual neural machine translate that enables zero-resource machine translation. When used together with novel many-to-one translation strategies, we empirically show that …

  23. Fully Character-Level Neural Machine Translation without Explicit Segmentation

    2017 · Transactions of the Association for Computational Linguistics

    Most existing machine translation systems operate at the level of words, relying on explicit segmentation to extract tokens. We introduce a neural machine translation (NMT) model that maps a source character sequence to a target …

  24. Learning to Parse and Translate Improves Neural Machine Translation

    2017

    There has been relatively little attention to incorporating linguistic prior to neural machine translation. Much of the previous work was further constrained to considering linguistic prior on the source side. In this paper, we propose …

  25. Nematus: a Toolkit for Neural Machine Translation

    2017

    Rico Sennrich, Orhan Firat, Kyunghyun Cho, Alexandra Birch, Barry Haddow, Julian Hitschler, Marcin Junczys-Dowmunt, Samuel Läubli, Antonio Valerio Miceli Barone, Jozef Mokry, Maria Nădejde. Proceedings of the Software Demonstrations of the 15th Conference of the …

  26. SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine

    2017 · arXiv (Cornell University)

    We publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect a full pipeline of …

  27. Search Engine Guided Neural Machine Translation

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    In this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. …

  28. Code-Switched Named Entity Recognition with Embedding Attention

    2018

    We describe our work for the CALCS 2018 shared task on named entity recognition on code-switched data. Our system ranked first place for MS Arabic-Egyptian named entity recognition and third place for English-Spanish.

  29. Meta-Learning for Low-Resource Neural Machine Translation

    2018

    In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT). We frame low-resource translation as a metalearning problem, and we learn …

  30. Passage Re-ranking with BERT

    2019 · arXiv (Cornell University)

    Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural …

  31. Non-Monotonic Sequential Text Generation

    2019 · arXiv (Cornell University)

    Standard sequential generation methods assume a pre-specified generation order, such as text generation methods which generate words from left to right. In this work, we propose a framework for training models of text generation that …

  32. Document Expansion by Query Prediction

    2019 · arXiv (Cornell University)

    One technique to improve the retrieval effectiveness of a search engine is to expand documents with terms that are related or representative of the documents' content.From the perspective of a question answering system, this might …

  33. Attention-Based Models for Speech Recognition

    2015 · arXiv (Cornell University)

    Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend …

  34. Zero-Shot Transfer Learning for Event Extraction

    2018

    Most previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort. We take a fresh look at event extraction …

  35. Task-Oriented Query Reformulation with Reinforcement Learning

    2017

    Search engines play an important role in our everyday lives by assisting us in finding the information we need. When we input a complex query, however, results are often far from satisfactory. In this work, …

  36. Dynamic Meta-Embeddings for Improved Sentence Representations

    2018

    While one of the first steps in many NLP systems is selecting what pre-trained word embeddings to use, we argue that such a step is better left for neural networks to figure out by themselves. …

  37. Neural Text Generation with Unlikelihood Training

    2019 · arXiv (Cornell University)

    Neural text generation is a key tool in natural language applications, but it is well known there are major problems at its core. In particular, standard likelihood training and decoding leads to dull and repetitive …

  38. Multi-Stage Document Ranking with BERT

    2019 · arXiv (Cornell University)

    The advent of deep neural networks pre-trained via language modeling tasks has spurred a number of successful applications in natural language processing. This work explores one such popular model, BERT, in the context of document …

  39. Insertion-based Decoding with Automatically Inferred Generation Order

    2019 · Transactions of the Association for Computational Linguistics

    Conventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. …

  40. Neural Machine Translation with Byte-Level Subwords

    2020 · Proceedings of the AAAI Conference on Artificial Intelligence

    Almost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up …

  41. Asking and Answering Questions to Evaluate the Factual Consistency of Summaries

    2020

    Practical applications of abstractive summarization models are limited by frequent factual inconsistencies with respect to their input. Existing automatic evaluation metrics for summarization are largely insensitive to such errors. We propose QAGS, 1 an automatic …

  42. True Few-Shot Learning with Language Models

    2021 · arXiv (Cornell University)

    Pretrained language models (LMs) perform well on many tasks even when learning from a few examples, but prior work uses many held-out examples to tune various aspects of learning, such as hyperparameters, training objectives, and …