Researcher profile

Anders Søgaard

13 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Reading metrics for estimating task efficiency with MT output

    2015

    We show that metrics derived from recording gaze while reading, are better proxies for machine translation quality than automated metrics. With reliable eyetracking technologies becoming available for home computers and mobile devices, such metrics are …

  2. Cross-lingual RST Discourse Parsing

    2017

    Discourse parsing is an integral part of understanding information flow and argumentative structure in documents. Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank. However, discourse treebanks for …

  3. Cross-lingual and cross-domain discourse segmentation of entire documents

    2017

    Discourse segmentation is a crucial step in building end-to-end discourse parsers. However, discourse segmenters only exist for a few languages and domains. Typically they only detect intra-sentential segment boundaries, assuming gold standard sentence and token …

  4. Does syntax help discourse segmentation? Not so much

    2017

    Discourse segmentation is the first step in building discourse parsers. Most work on discourse segmentation does not scale to real-world discourse parsing across languages, for two reasons: (i) models rely on constituent trees, and (ii) …

  5. Model-based annotation of coreference

    2019 · arXiv (Cornell University)

    Humans do not make inferences over texts, but over models of what texts are about. When annotators are asked to annotate coreferent spans of text, it is therefore a somewhat unnatural task. This paper presents …

  6. Common Sense Bias in Semantic Role Labeling

    2021

    Large-scale language models such as ELMo and BERT have pushed the horizon of what is possible in semantic role labeling (SRL), solving the out-of-vocabulary problem and enabling end-to-end systems, but they have also introduced significant …

  7. Multi hash embeddings in spaCy

    2022 · arXiv (Cornell University)

    The distributed representation of symbols is one of the key technologies in machine learning systems today, playing a pivotal role in modern natural language processing. Traditional word embeddings associate a separate vector with each word. …

  8. The Sensitivity of Annotator Bias to Task Definitions in Argument Mining

    2022 · IT University Of Copenhagen (IT University of Copenhagen)

    NLP models are dependent on the data they are trained on, including how this data is annotated. NLP research increasingly examines the social biases of models, but often in the light of their training data …

  9. A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments

    2017

    While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set (sentence IDs) …

  10. Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm

    2017

    NLP tasks are often limited by scarcity of manually annotated data. In social media sentiment analysis and related tasks, researchers have therefore used binarized emoticons and specific hashtags as forms of distant supervision. Our paper …

  11. A Survey of Cross-lingual Word Embedding Models

    2019 · Journal of Artificial Intelligence Research

    Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …

  12. Unsupervised Cross-Lingual Representation Learning

    2019

    In this tutorial, we provide a comprehensive survey of the exciting recent work on cutting-edge weakly-supervised and unsupervised cross-lingual word representations. After providing a brief history of supervised cross-lingual word representations, we focus on: 1) …

  13. A Survey of Cross-lingual Word Embedding Models

    2018 · Apollo (University of Cambridge)

    Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …