Jun Xu
13 papers in the PaperMetrix corpus
Papers by this author
-
Clinical Named Entity Recognition Using Deep Learning Models.
2017 · PubMed
Clinical Named Entity Recognition (NER) is a critical natural language processing (NLP) task to extract important concepts (named entities) from clinical narratives. Researchers have extensively investigated machine learning models for clinical NER. Recently, there have …
-
A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech
2021 · arXiv (Cornell University)
In this paper, we present a neural model for joint dropped pronoun recovery (DPR) and conversational discourse parsing (CDP) in Chinese conversational speech. We show that DPR and CDP are closely related, and a joint …
-
A Brief History of Recommender Systems
2022 · arXiv (Cornell University)
Soon after the invention of the Internet, the recommender system emerged and related technologies have been extensively studied and applied by both academia and industry. Currently, recommender system has become one of the most successful …
-
YuLan: An Open-source Large Language Model
2024 · arXiv (Cornell University)
Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training …
-
Response probability distribution estimation of expensive computer simulators: A Bayesian active learning perspective using Gaussian process regression
2024 · arXiv (Cornell University)
Estimation of the response probability distributions of computer simulators in the presence of randomness is a crucial task in many fields. However, achieving this task with guaranteed accuracy remains an open computational challenge, especially for …
-
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
2025 · arXiv (Cornell University)
High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus …
-
Trigger3:Refining Query Correction via Adaptive Model Selector
2025 · Proceedings of the AAAI Conference on Artificial Intelligence
In search scenarios, user experience can be hindered by erroneous queries due to typos, voice errors, or knowledge gaps. Therefore, query correction is crucial for search engines. Current correction models, usually small models trained on …
-
Learning Hierarchical Representation Model for NextBasket Recommendation
2015
Next basket recommendation is a crucial task in market basket analysis. Given a user's purchase history, usually a sequence of transaction data, one attempts to build a recommender that can predict the next few items …
-
Clinical Abbreviation Disambiguation Using Neural Word Embeddings
2015
This study examined the use of neural word embeddings for clinical abbreviation disambiguation, a special case of word sense disambiguation (WSD). We investigated three different methods for deriving word embeddings from a large unlabeled clinical …
-
A Deep Architecture for Semantic Matching with Multiple Positional Sentence Representations
2016 · Proceedings of the AAAI Conference on Artificial Intelligence
Matching natural language sentences is central for many applications such as information retrieval and question answering. Existing deep models rely on a single sentence representation or multiple granularity representations for matching. However, such methods cannot …
-
Modeling Document Novelty with Neural Tensor Network for Search Result Diversification
2016
Search result diversification has attracted considerable attention as a means to tackle the ambiguous or multi-faceted information needs of users. One of the key problems in search result diversification is novelty, that is, how to …
-
Reinforcement Learning to Rank with Markov Decision Process
2017
One of the central issues in learning to rank for information retrieval is to develop algorithms that construct ranking models by directly optimizing evaluation measures such as normalized discounted cumulative gain~(ND CG). Existing methods usually …
-
A Deep Architecture for Semantic Matching with Multiple Positional Sentence Representations
2015 · arXiv (Cornell University)
Matching natural language sentences is central for many applications such as information retrieval and question answering. Existing deep models rely on a single sentence representation or multiple granularity representations for matching. However, such methods cannot …