Donald Metzler
7 papers in the PaperMetrix corpus
Papers by this author
-
Learning to Rank with Selection Bias in Personal Search
2016
Click-through data has proven to be a critical resource for improving search ranking quality. Though a large amount of click data can be easily collected by search engines, various biases make it difficult to fully …
-
Multi-Task Learning for Email Search Ranking with Auxiliary Query Clustering
2018
User information needs vary significantly across different tasks, and therefore their queries will also differ considerably in their expressiveness and semantics. Many studies have been proposed to model such query diversity by obtaining query types …
-
Multitask Mixture of Sequential Experts for User Activity Streams
2020
It is often desirable to model multiple objectives in real-world web applications, such as user satisfaction and user engagement in recommender systems. Multi-task learning has become the standard approach for such applications recently.
-
Emergent Abilities of Large Language Models
2022 · arXiv (Cornell University)
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …
-
Confident Adaptive Language Modeling
2022 · arXiv (Cornell University)
Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in the models' size, potentially leading to slow and costly use …
-
Position Bias Estimation for Unbiased Learning to Rank in Personal Search
2018
A well-known challenge in learning from click data is its inherent bias and most notably position bias. Traditional click models aim to extract the ‹query, document› relevance and the estimated bias is usually discarded after …
-
UL2: Unifying Language Learning Paradigms
2022 · arXiv (Cornell University)
Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …