Researcher profile

Sebastian Ruder

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Data Selection Strategies for Multi-Domain Sentiment Analysis

    2017 · arXiv (Cornell University)

    Domain adaptation is important in sentiment analysis as sentiment-indicating words vary between domains. Recently, multi-domain adaptation has become more pervasive, but existing approaches train on all available source domains including dissimilar ones. However, the selection …

  2. Towards a continuous modeling of natural language domains

    2016 · arXiv (Cornell University)

    Humans continuously adapt their style and language to a variety of domains. However, a reliable definition of `domain' has eluded researchers thus far. Additionally, the notion of discrete domains stands in contrast to the multiplicity …

  3. Mind the Gap: Assessing Temporal Generalization in Neural Language Models

    2021 · arXiv (Cornell University)

    Our world is open-ended, non-stationary, and constantly evolving; thus what we talk about and how we talk about it change over time. This inherent dynamic nature of language contrasts with the current static language modelling …

  4. NL-Augmenter 🦎 → 🐍 A Framework for Task-Sensitive Natural Language Augmentation

    2023 · Northern European Journal of Language Technology

    Data augmentation is an important method for evaluating the robustness of and enhancing the diversity of training data for natural language processing (NLP) models. In this paper, we present NL-Augmenter, a new participatory Python-based natural …

  5. One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia

    2022 · arXiv (Cornell University)

    NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth …

  6. NusaCrowd: Open Source Initiative for Indonesian NLP Resources

    2022 · arXiv (Cornell University)

    We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data …

  7. A Survey of Cross-lingual Word Embedding Models

    2019 · Journal of Artificial Intelligence Research

    Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …

  8. Transfer Learning in Natural Language Processing

    2019

    The classic supervised machine learning paradigm is based on learning in isolation, a single predictive model for a task using a single dataset. This approach requires a large number of training examples and performs best …

  9. A Hierarchical Multi-Task Approach for Learning Embeddings from Semantic Tasks

    2019

    Much effort has been devoted to evaluate whether multi-task learning can be leveraged to learn rich representations that can be used in various Natural Language Processing (NLP) down-stream applications. However, there is still a lack …

  10. Unsupervised Cross-Lingual Representation Learning

    2019

    In this tutorial, we provide a comprehensive survey of the exciting recent work on cutting-edge weakly-supervised and unsupervised cross-lingual word representations. After providing a brief history of supervised cross-lingual word representations, we focus on: 1) …

  11. On the Cross-lingual Transferability of Monolingual Representations

    2020

    State-of-the-art unsupervised multilingual models (e.g., multilingual BERT) have been shown to generalize in a zero-shot cross-lingual setting. This generalization ability has been attributed to the use of a shared subword vocabulary and joint training across …

  12. XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

    2020 · arXiv (Cornell University)

    Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, …

  13. A Survey of Cross-lingual Word Embedding Models

    2018 · Apollo (University of Cambridge)

    Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …

  14. MasakhaNER: Named Entity Recognition for African Languages

    2021 · Transactions of the Association for Computational Linguistics

    Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …