Ali Farhadi
6 papers in the PaperMetrix corpus
Papers by this author
-
2 OLMo 2 Furious
2024 · arXiv (Cornell University)
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model …
-
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
2025 · arXiv (Cornell University)
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output …
-
Bidirectional Attention Flow for Machine Comprehension
2016 · arXiv (Cornell University)
Machine comprehension (MC), answering a query about a given context paragraph, requires modeling complex interactions between the context and the query. Recently, attention mechanisms have been successfully extended to MC. Typically these methods use attention …
-
HellaSwag: Can a Machine Really Finish Your Sentence?
2019
introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the …
-
Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index
2019
Existing open-domain question answering (QA) models are not suitable for real-time usage because they need to process several long documents on-demand for every input query, which is computationally prohibitive. In this paper, we introduce query-agnostic …
-
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
2020 · arXiv (Cornell University)
Fine-tuning pretrained contextual word embedding models to supervised downstream tasks has become commonplace in natural language processing. This process, however, is often brittle: even with the same hyperparameter values, distinct random seeds can lead to …