Researcher profile

Jimmy Lin

29 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. JavaScript Convolutional Neural Networks for Keyword Spotting in the Browser: An Experimental Analysis

    2018 · arXiv (Cornell University)

    Used for simple commands recognition on devices from smart routers to mobile phones, keyword spotting systems are everywhere. Ubiquitous as well are web applications, which have grown in popularity and complexity over the last decade …

  2. Investigating the Limitations of Transformers with Simple Arithmetic Tasks

    2021 · arXiv (Cornell University)

    The ability to perform arithmetic tasks is a remarkable trait of human intelligence and might form a critical component of more complex reasoning tasks. In this work, we investigate if the surface form of a …

  3. Overview of the TREC 2021 deep learning track

    2025 · arXiv (Cornell University)

    This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels available for both passage and …

  4. Ten best practices to strengthen stewardship and sharing of water science data in Canada

    2021 · Hydrological Processes

    Abstract Water science data are a valuable asset that both underpins the original research project and bolsters new research questions, particularly in view of the increasingly complex water issues facing Canada and the world. Whilst …

  5. Improving Query Representations for Dense Retrieval with Pseudo Relevance Feedback: A Reproducibility Study

    2021 · arXiv (Cornell University)

    Pseudo-Relevance Feedback (PRF) utilises the relevance signals from the top-k passages from the first round of retrieval to perform a second round of retrieval aiming to improve search effectiveness. A recent research direction has been …

  6. “Low-Resource” Text Classification: A Parameter-Free Classification Method with Compressors

    2023

    Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them …

  7. Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking

    2022

    In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical …

  8. Approximating Human-Like Few-shot Learning with GPT-based Compression

    2023 · arXiv (Cornell University)

    In this work, we conceptualize the learning process as information compression. We seek to equip generative pre-trained models with human-like learning capabilities that enable data compression during inference. We present a novel approach that utilizes …

  9. RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models

    2023 · arXiv (Cornell University)

    Researchers have successfully applied large language models (LLMs) such as ChatGPT to reranking in an information retrieval context, but to date, such work has mostly been built on proprietary models hidden behind opaque API endpoints. …

  10. BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

    2025 · arXiv (Cornell University)

    Deep-Research agents, which integrate large language models (LLMs) with search tools, have shown success in improving the effectiveness of handling complex queries that require iterative search planning and reasoning over search results. Evaluations on current …

  11. Multi-Perspective Sentence Similarity Modeling with Convolutional Neural Networks

    2015

    Modeling sentence similarity is complicated by the ambiguity and variability of linguistic expression. To cope with these challenges, we propose a model for comparing sentences that uses a multiplicity of perspectives. We first model each …

  12. Pairwise Word Interaction Modeling with Deep Neural Networks for Semantic Similarity Measurement

    2016

    Textual similarity measurement is a challenging problem, as it requires understanding the semantics of input sentences. Most previous neural network models use coarse-grained sentence modeling, which has difficulty capturing fine-grained word-level information for semantic comparisons. …

  13. Noise-Contrastive Estimation for Answer Selection with Deep Neural Networks

    2016

    We study answer selection for question answering, in which given a question and a set of candidate answer sentences, the goal is to identify the subset that contains the answer. Unlike previous work which treats …

  14. Anserini

    2017

    Software toolkits play an essential role in information retrieval research. Most open-source toolkits developed by academics are designed to facilitate the evaluation of retrieval models over standard test collections. Efforts are generally directed toward better …

  15. Anserini

    2018 · Journal of Data and Information Quality

    This work tackles the perennial problem of reproducible baselines in information retrieval research, focusing on bag-of-words ranking models. Although academic information retrieval researchers have a long history of building and sharing systems, they are primarily …

  16. End-to-End Open-Domain Question Answering with

    2019

    Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, Jimmy Lin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.

  17. Models and Data for Simple Applications of BERT for Ad Hoc Document Retrieval

    2019 · arXiv (Cornell University)

    Following recent successes in applying BERT to question answering, we explore simple applications to ad hoc document retrieval. This required confronting the challenge posed by documents that are typically longer than the length of input …

  18. Distilling Task-Specific Knowledge from BERT into Simple Neural Networks

    2019 · arXiv (Cornell University)

    In the natural language processing literature, neural networks are becoming increasingly deeper and complex. The recent poster child of this trend is the deep language representation model, which includes BERT, ELMo, and GPT. These developments …

  19. Simple BERT Models for Relation Extraction and Semantic Role Labeling

    2019 · arXiv (Cornell University)

    We present simple BERT-based models for relation extraction and semantic role labeling. In recent years, state-of-the-art performance has been achieved using neural models by incorporating lexical and syntactic features such as part-of-speech tags and dependency …

  20. Document Expansion by Query Prediction

    2019 · arXiv (Cornell University)

    One technique to improve the retrieval effectiveness of a search engine is to expand documents with terms that are related or representative of the documents' content.From the perspective of a question answering system, this might …

  21. Data Augmentation for BERT Fine-Tuning in Open-Domain Question Answering

    2019 · arXiv (Cornell University)

    Recently, a simple combination of passage retrieval using off-the-shelf IR techniques and a BERT reader was found to be very effective for question answering directly on Wikipedia, yielding a large improvement over the previous state …

  22. DocBERT: BERT for Document Classification

    2019 · arXiv (Cornell University)

    We present, to our knowledge, the first application of BERT to document classification. A few characteristics of the task might lead one to think that BERT is not the most appropriate model: syntactic structures matter …

  23. Strong Baselines for Simple Question Answering over Knowledge Graphs with and without Neural Networks

    2018

    Salman Mohammed, Peng Shi, Jimmy Lin. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 2018.

  24. Aligning Cross-Lingual Entities with Multi-Aspect Information

    2019

    Hsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). …

  25. Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval

    2019

    Zeynep Akkalyoncu Yilmaz, Wei Yang, Haotian Zhang, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  26. Multi-Stage Document Ranking with BERT

    2019 · arXiv (Cornell University)

    The advent of deep neural networks pre-trained via language modeling tasks has spurred a number of successful applications in natural language processing. This work explores one such popular model, BERT, in the context of document …

  27. DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference

    2020

    Large-scale pre-trained language models such as BERT have brought significant improvements to NLP applications. However, they are also notorious for being slow in inference, which makes them difficult to deploy in realtime applications. We propose …

  28. Pretrained Transformers for Text Ranking: BERT and Beyond

    2020 · arXiv (Cornell University)

    The goal of text ranking is to generate an ordered list of texts retrieved from a corpus in response to a query. Although the most common formulation of text ranking is search, instances of the …

  29. Document Ranking with a Pretrained Sequence-to-Sequence Model

    2020

    This work proposes the use of a pretrained sequence-to-sequence model for document ranking. Our approach is fundamentally different from a commonly adopted classificationbased formulation based on encoder-only pretrained transformer architectures such as BERT. We show …