Mostafa Dehghani
8 papers in the PaperMetrix corpus
Papers by this author
-
Building a multi-domain comparable corpus using a learning to rank method
2016 · Natural Language Engineering
Abstract Comparable corpora are key translation resources for both languages and domains with limited linguistic resources. The existing approaches for building comparable corpora are mostly based on ranking candidate documents in the target language for …
-
Avoiding Your Teacher's Mistakes: Training Neural Networks with Controlled Weak Supervision
2017 · arXiv (Cornell University)
Training deep neural networks requires massive amounts of training data, but for many tasks only limited labeled data is available. This makes weak supervision attractive, using weak or noisy signals like the output of heuristic …
-
Confident Adaptive Language Modeling
2022 · arXiv (Cornell University)
Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in the models' size, potentially leading to slow and costly use …
-
Neural Ranking Models with Weak Supervision
2017
Despite the impressive improvements achieved by unsupervised deep neural networks in computer vision and NLP tasks, such improvements have not yet been observed in ranking for information retrieval. The reason may be the complexity of …
-
Universal Transformers
2018 · arXiv (Cornell University)
Recurrent neural networks (RNNs) sequentially process data by updating their state with each new data point, and have long been the de facto choice for sequence modeling tasks. However, their inherently sequential computation makes them …
-
From Neural Re-Ranking to Neural Ranking
2018
The availability of massive data and computing power allowing for effective data driven neural approaches is having a major impact on machine learning and information retrieval research, but these models have a basic problem with …
-
UL2: Unifying Language Learning Paradigms
2022 · arXiv (Cornell University)
Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …
-
Scaling Instruction-Finetuned Language Models
2022 · arXiv (Cornell University)
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …