Researcher profile

Mostafa Dehghani

8 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Building a multi-domain comparable corpus using a learning to rank method

    2016 · Natural Language Engineering

    Abstract Comparable corpora are key translation resources for both languages and domains with limited linguistic resources. The existing approaches for building comparable corpora are mostly based on ranking candidate documents in the target language for …

  2. Avoiding Your Teacher's Mistakes: Training Neural Networks with Controlled Weak Supervision

    2017 · arXiv (Cornell University)

    Training deep neural networks requires massive amounts of training data, but for many tasks only limited labeled data is available. This makes weak supervision attractive, using weak or noisy signals like the output of heuristic …

  3. Confident Adaptive Language Modeling

    2022 · arXiv (Cornell University)

    Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in the models' size, potentially leading to slow and costly use …

  4. Neural Ranking Models with Weak Supervision

    2017

    Despite the impressive improvements achieved by unsupervised deep neural networks in computer vision and NLP tasks, such improvements have not yet been observed in ranking for information retrieval. The reason may be the complexity of …

  5. Universal Transformers

    2018 · arXiv (Cornell University)

    Recurrent neural networks (RNNs) sequentially process data by updating their state with each new data point, and have long been the de facto choice for sequence modeling tasks. However, their inherently sequential computation makes them …

  6. From Neural Re-Ranking to Neural Ranking

    2018

    The availability of massive data and computing power allowing for effective data driven neural approaches is having a major impact on machine learning and information retrieval research, but these models have a basic problem with …

  7. UL2: Unifying Language Learning Paradigms

    2022 · arXiv (Cornell University)

    Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …

  8. Scaling Instruction-Finetuned Language Models

    2022 · arXiv (Cornell University)

    Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …