Jimmy Lin
29 papers in the PaperMetrix corpus
Papers by this author
-
JavaScript Convolutional Neural Networks for Keyword Spotting in the Browser: An Experimental Analysis
2018 · arXiv (Cornell University)
Used for simple commands recognition on devices from smart routers to mobile phones, keyword spotting systems are everywhere. Ubiquitous as well are web applications, which have grown in popularity and complexity over the last decade …
-
Investigating the Limitations of Transformers with Simple Arithmetic Tasks
2021 · arXiv (Cornell University)
The ability to perform arithmetic tasks is a remarkable trait of human intelligence and might form a critical component of more complex reasoning tasks. In this work, we investigate if the surface form of a …
-
Overview of the TREC 2021 deep learning track
2025 · arXiv (Cornell University)
This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels available for both passage and …
-
Ten best practices to strengthen stewardship and sharing of water science data in Canada
2021 · Hydrological Processes
Abstract Water science data are a valuable asset that both underpins the original research project and bolsters new research questions, particularly in view of the increasingly complex water issues facing Canada and the world. Whilst …
-
Improving Query Representations for Dense Retrieval with Pseudo Relevance Feedback: A Reproducibility Study
2021 · arXiv (Cornell University)
Pseudo-Relevance Feedback (PRF) utilises the relevance signals from the top-k passages from the first round of retrieval to perform a second round of retrieval aiming to improve search effectiveness. A recent research direction has been …
-
“Low-Resource” Text Classification: A Parameter-Free Classification Method with Compressors
2023
Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them …
-
Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking
2022
In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical …
-
Approximating Human-Like Few-shot Learning with GPT-based Compression
2023 · arXiv (Cornell University)
In this work, we conceptualize the learning process as information compression. We seek to equip generative pre-trained models with human-like learning capabilities that enable data compression during inference. We present a novel approach that utilizes …
-
RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models
2023 · arXiv (Cornell University)
Researchers have successfully applied large language models (LLMs) such as ChatGPT to reranking in an information retrieval context, but to date, such work has mostly been built on proprietary models hidden behind opaque API endpoints. …
-
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
2025 · arXiv (Cornell University)
Deep-Research agents, which integrate large language models (LLMs) with search tools, have shown success in improving the effectiveness of handling complex queries that require iterative search planning and reasoning over search results. Evaluations on current …
-
Multi-Perspective Sentence Similarity Modeling with Convolutional Neural Networks
2015
Modeling sentence similarity is complicated by the ambiguity and variability of linguistic expression. To cope with these challenges, we propose a model for comparing sentences that uses a multiplicity of perspectives. We first model each …
-
Pairwise Word Interaction Modeling with Deep Neural Networks for Semantic Similarity Measurement
2016
Textual similarity measurement is a challenging problem, as it requires understanding the semantics of input sentences. Most previous neural network models use coarse-grained sentence modeling, which has difficulty capturing fine-grained word-level information for semantic comparisons. …
-
Noise-Contrastive Estimation for Answer Selection with Deep Neural Networks
2016
We study answer selection for question answering, in which given a question and a set of candidate answer sentences, the goal is to identify the subset that contains the answer. Unlike previous work which treats …
-
Anserini
2017
Software toolkits play an essential role in information retrieval research. Most open-source toolkits developed by academics are designed to facilitate the evaluation of retrieval models over standard test collections. Efforts are generally directed toward better …
-
Anserini
2018 · Journal of Data and Information Quality
This work tackles the perennial problem of reproducible baselines in information retrieval research, focusing on bag-of-words ranking models. Although academic information retrieval researchers have a long history of building and sharing systems, they are primarily …
-
End-to-End Open-Domain Question Answering with
2019
Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, Jimmy Lin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.
-
Models and Data for Simple Applications of BERT for Ad Hoc Document Retrieval
2019 · arXiv (Cornell University)
Following recent successes in applying BERT to question answering, we explore simple applications to ad hoc document retrieval. This required confronting the challenge posed by documents that are typically longer than the length of input …
-
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
2019 · arXiv (Cornell University)
In the natural language processing literature, neural networks are becoming increasingly deeper and complex. The recent poster child of this trend is the deep language representation model, which includes BERT, ELMo, and GPT. These developments …
-
Simple BERT Models for Relation Extraction and Semantic Role Labeling
2019 · arXiv (Cornell University)
We present simple BERT-based models for relation extraction and semantic role labeling. In recent years, state-of-the-art performance has been achieved using neural models by incorporating lexical and syntactic features such as part-of-speech tags and dependency …
-
Document Expansion by Query Prediction
2019 · arXiv (Cornell University)
One technique to improve the retrieval effectiveness of a search engine is to expand documents with terms that are related or representative of the documents' content.From the perspective of a question answering system, this might …
-
Data Augmentation for BERT Fine-Tuning in Open-Domain Question Answering
2019 · arXiv (Cornell University)
Recently, a simple combination of passage retrieval using off-the-shelf IR techniques and a BERT reader was found to be very effective for question answering directly on Wikipedia, yielding a large improvement over the previous state …
-
DocBERT: BERT for Document Classification
2019 · arXiv (Cornell University)
We present, to our knowledge, the first application of BERT to document classification. A few characteristics of the task might lead one to think that BERT is not the most appropriate model: syntactic structures matter …
-
Strong Baselines for Simple Question Answering over Knowledge Graphs with and without Neural Networks
2018
Salman Mohammed, Peng Shi, Jimmy Lin. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 2018.
-
Aligning Cross-Lingual Entities with Multi-Aspect Information
2019
Hsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). …
-
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval
2019
Zeynep Akkalyoncu Yilmaz, Wei Yang, Haotian Zhang, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Multi-Stage Document Ranking with BERT
2019 · arXiv (Cornell University)
The advent of deep neural networks pre-trained via language modeling tasks has spurred a number of successful applications in natural language processing. This work explores one such popular model, BERT, in the context of document …
-
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
2020
Large-scale pre-trained language models such as BERT have brought significant improvements to NLP applications. However, they are also notorious for being slow in inference, which makes them difficult to deploy in realtime applications. We propose …
-
Pretrained Transformers for Text Ranking: BERT and Beyond
2020 · arXiv (Cornell University)
The goal of text ranking is to generate an ordered list of texts retrieved from a corpus in response to a query. Although the most common formulation of text ranking is search, instances of the …
-
Document Ranking with a Pretrained Sequence-to-Sequence Model
2020
This work proposes the use of a pretrained sequence-to-sequence model for document ranking. Our approach is fundamentally different from a commonly adopted classificationbased formulation based on encoder-only pretrained transformer architectures such as BERT. We show …