Luca Soldaini
4 papers in the PaperMetrix corpus
Papers by this author
-
One-Shot Labeling for Automatic Relevance Estimation
2023
Dealing with unjudged documents ("holes") in relevance assessments is a perennial problem when evaluating search systems with offline experiments. Holes can reduce the apparent effectiveness of retrieval systems during evaluation and introduce biases in models …
-
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
2024 · arXiv (Cornell University)
We present SciRIFF (Scientific Resource for Instruction-Following and Finetuning), a dataset of 137K instruction-following instances for training and evaluation, covering 54 tasks. These tasks span five core scientific literature understanding capabilities: information extraction, summarization, question …
-
2 OLMo 2 Furious
2024 · arXiv (Cornell University)
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model …
-
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
2025 · arXiv (Cornell University)
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output …