Caiming Xiong
29 papers in the PaperMetrix corpus
Papers by this author
-
A Deep Reinforced Model for Abstractive Summarization
2017 · arXiv (Cornell University)
Attentional, RNN-based encoder-decoder models for abstractive summarization have achieved good performance on short input and output sequences. For longer documents and summaries however these models often include repetitive and incoherent phrases. We introduce a neural …
-
Efficient and Robust Question Answering from Minimal Context over Documents
2018 · ArXiv.org
Neural models for question answering (QA) over documents have achieved significant performance improvements. Although effective, these models do not scale to large corpora due to their complex modeling of interactions between the document and the …
-
Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering
2019 · International Conference on Learning Representations
End-to-end neural models have made significant progress in question answering, however recent studies show that these models implicitly assume that the answer and evidence appear close together in a single document. In this work, we …
-
Unsupervised Out-of-Domain Detection via Pre-trained Transformers
2021 · arXiv (Cornell University)
Deployed real-world machine learning applications are often subject to uncontrolled and even potentially malicious inputs. Those out-of-domain inputs can lead to unpredictable outputs and sometimes catastrophic safety issues. Prior studies on out-of-domain detection require in-domain …
-
Don’t Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification
2021 · International Conference on Machine Learning
Modern machine learning models with high accuracy are often miscalibrated -- the predicted top probability does not reflect the actual accuracy, and tends to be over-confident. It is commonly believed that such over-confidence is mainly …
-
Modeling Multi-hop Question Answering as Single Sequence Prediction
2022 · arXiv (Cornell University)
Fusion-in-decoder (Fid) (Izacard and Grave, 2020) is a generative question answering (QA) model that leverages passage retrieval with a pre-trained transformer and pushed the state of the art on single-hop QA. However, the complexity of …
-
Use All The Labels: A Hierarchical Multi-Label Contrastive Learning Framework
2022 · 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Current contrastive learning frameworks focus on leveraging a single supervisory signal to learn representations, which limits the efficacy on unseen data and downstream tasks. In this paper, we present a hierarchical multi-label representation learning framework …
-
Improved Online Conformal Prediction via Strongly Adaptive Online Learning
2023 · arXiv (Cornell University)
We study the problem of uncertainty quantification via prediction sets, in an online setting where the data distribution may vary arbitrarily over time. Recent work develops online conformal prediction techniques that leverage regret minimization algorithms …
-
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
2023 · arXiv (Cornell University)
Upon its release in late 2022, ChatGPT has brought a seismic shift in the entire landscape of AI, both in research and commerce. Through instruction-tuning a large language model (LLM) with supervised fine-tuning and reinforcement …
-
Personalized Multi-task Training for Recommender System
2024 · arXiv (Cornell University)
In the vast landscape of internet information, recommender systems (RecSys) have become essential for guiding users through a sea of choices aligned with their preferences. These systems have applications in diverse domains, such as news …
-
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
2025
Jierui Li, Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: …
-
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
2025 · arXiv (Cornell University)
Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified …
-
Pointer Sentinel Mixture Models
2016 · arXiv (Cornell University)
Recent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies. Even then they struggle to predict rare or unseen words even …
-
Dynamic Coattention Networks For Question Answering
2016 · arXiv (Cornell University)
Several deep learning models have been proposed for question answering. However, due to their single-pass nature, they have no way to recover from local maxima corresponding to incorrect answers. To address this problem, we introduce …
-
A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
2017
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks. Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in …
-
Learned in Translation: Contextualized Word Vectors
2017 · arXiv (Cornell University)
Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with …
-
DCN+: Mixed Objective and Deep Residual Coattention for Question Answering
2017 · arXiv (Cornell University)
Traditional models for question answering optimize using cross entropy loss, which encourages exact answers at the cost of penalizing nearby or overlapping answers that are sometimes equally accurate. We propose a mixed objective that combines …
-
Non-Autoregressive Neural Machine Translation
2017 · arXiv (Cornell University)
Existing approaches to neural machine translation condition each output word on previously generated outputs. We introduce a model that avoids this autoregressive property and produces its outputs in parallel, allowing an order of magnitude lower …
-
The Natural Language Decathlon: Multitask Learning as Question Answering
2018 · arXiv (Cornell University)
Deep learning has improved performance on many natural language processing (NLP) tasks individually. However, general NLP models cannot emerge within a paradigm that focuses on the particularities of a single metric, dataset, and task. We …
-
XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
2019 · arXiv (Cornell University)
While natural language processing systems often focus on a single language, multilingual transfer learning has the potential to improve performance, especially for low-resource languages. We introduce XLDA, cross-lingual data augmentation, a method that replaces a …
-
A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
2016 · arXiv (Cornell University)
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks. Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in …
-
Quasi-Recurrent Neural Networks
2016 · arXiv (Cornell University)
Recurrent neural networks are a powerful tool for modeling sequential data, but the dependence of each timestep's computation on the previous timestep's output limits parallelism and makes RNNs unwieldy for very long sequences. We introduce …
-
CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
2019
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao …
-
CTRL: A Conditional Transformer Language Model for Controllable Generation
2019 · arXiv (Cornell University)
Large-scale language models show promising text generation capabilities, but users cannot easily control particular aspects of the generated text. We release CTRL, a 1.63 billion-parameter conditional transformer language model, trained to condition on control codes …
-
BERT is Not an Interlingua and the Bias of Tokenization
2019
Multilingual transfer learning can benefit both high-and low-resource languages, but the source of these improvements is not well understood. Cananical Correlation Analysis (CCA) of the internal representations of a pretrained, multilingual BERT model reveals that …
-
Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
2019 · arXiv (Cornell University)
Answering questions that require multi-hop reasoning at web-scale necessitates retrieving multiple evidence documents, one of which often has little lexical or semantic relationship to the question. This paper introduces a new graph-based recurrent retrieval approach …
-
RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering
2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Existing KBQA approaches, despite achieving strong performance on i.i.d. test data, often struggle in generalizing to questions involving unseen KB schema items. Prior rankingbased approaches have shown some success in generalization, but suffer from the …
-
Intent Contrastive Learning for Sequential Recommendation
2022 · Proceedings of the ACM Web Conference 2022
Users’ interactions with items are driven by various intents (e.g., preparing for holiday gifts, shopping for fishing equipment, etc.). However, users’ underlying intents are often unobserved/latent, making it challenging to leverage such latent intents for …
-
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
2022
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir …