Researcher profile

Kyle Lo

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. SciBERT: Pretrained Contextualized Embeddings for Scientific Text

    2019

    Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, …

  2. SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature

    2024 · arXiv (Cornell University)

    We present SciRIFF (Scientific Resource for Instruction-Following and Finetuning), a dataset of 137K instruction-following instances for training and evaluation, covering 54 tasks. These tasks span five core scientific literature understanding capabilities: information extraction, summarization, question …

  3. 2 OLMo 2 Furious

    2024 · arXiv (Cornell University)

    We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model …

  4. Construction of the Literature Graph in Semantic Scholar

    2018

    Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, Rodney Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler Murray, Hsu-Han Ooi, Matthew Peters, Joanna Power, Sam …

  5. SciBERT: A Pretrained Language Model for Scientific Text

    2019

    Iz Beltagy, Kyle Lo, Arman Cohan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  6. S2ORC: The Semantic Scholar Open Research Corpus

    2020

    We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M …

  7. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    2022 · arXiv (Cornell University)

    Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …