Researcher profile

Jun Xu

13 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Clinical Named Entity Recognition Using Deep Learning Models.

    2017 · PubMed

    Clinical Named Entity Recognition (NER) is a critical natural language processing (NLP) task to extract important concepts (named entities) from clinical narratives. Researchers have extensively investigated machine learning models for clinical NER. Recently, there have …

  2. A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech

    2021 · arXiv (Cornell University)

    In this paper, we present a neural model for joint dropped pronoun recovery (DPR) and conversational discourse parsing (CDP) in Chinese conversational speech. We show that DPR and CDP are closely related, and a joint …

  3. A Brief History of Recommender Systems

    2022 · arXiv (Cornell University)

    Soon after the invention of the Internet, the recommender system emerged and related technologies have been extensively studied and applied by both academia and industry. Currently, recommender system has become one of the most successful …

  4. YuLan: An Open-source Large Language Model

    2024 · arXiv (Cornell University)

    Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training …

  5. Response probability distribution estimation of expensive computer simulators: A Bayesian active learning perspective using Gaussian process regression

    2024 · arXiv (Cornell University)

    Estimation of the response probability distributions of computer simulators in the presence of randomness is a crucial task in many fields. However, achieving this task with guaranteed accuracy remains an open computational challenge, especially for …

  6. Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

    2025 · arXiv (Cornell University)

    High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus …

  7. Trigger3:Refining Query Correction via Adaptive Model Selector

    2025 · Proceedings of the AAAI Conference on Artificial Intelligence

    In search scenarios, user experience can be hindered by erroneous queries due to typos, voice errors, or knowledge gaps. Therefore, query correction is crucial for search engines. Current correction models, usually small models trained on …

  8. Learning Hierarchical Representation Model for NextBasket Recommendation

    2015

    Next basket recommendation is a crucial task in market basket analysis. Given a user's purchase history, usually a sequence of transaction data, one attempts to build a recommender that can predict the next few items …

  9. Clinical Abbreviation Disambiguation Using Neural Word Embeddings

    2015

    This study examined the use of neural word embeddings for clinical abbreviation disambiguation, a special case of word sense disambiguation (WSD). We investigated three different methods for deriving word embeddings from a large unlabeled clinical …

  10. A Deep Architecture for Semantic Matching with Multiple Positional Sentence Representations

    2016 · Proceedings of the AAAI Conference on Artificial Intelligence

    Matching natural language sentences is central for many applications such as information retrieval and question answering. Existing deep models rely on a single sentence representation or multiple granularity representations for matching. However, such methods cannot …

  11. Modeling Document Novelty with Neural Tensor Network for Search Result Diversification

    2016

    Search result diversification has attracted considerable attention as a means to tackle the ambiguous or multi-faceted information needs of users. One of the key problems in search result diversification is novelty, that is, how to …

  12. Reinforcement Learning to Rank with Markov Decision Process

    2017

    One of the central issues in learning to rank for information retrieval is to develop algorithms that construct ranking models by directly optimizing evaluation measures such as normalized discounted cumulative gain~(ND CG). Existing methods usually …

  13. A Deep Architecture for Semantic Matching with Multiple Positional Sentence Representations

    2015 · arXiv (Cornell University)

    Matching natural language sentences is central for many applications such as information retrieval and question answering. Existing deep models rely on a single sentence representation or multiple granularity representations for matching. However, such methods cannot …