Anders Søgaard
13 papers in the PaperMetrix corpus
Papers by this author
-
Reading metrics for estimating task efficiency with MT output
2015
We show that metrics derived from recording gaze while reading, are better proxies for machine translation quality than automated metrics. With reliable eyetracking technologies becoming available for home computers and mobile devices, such metrics are …
-
Cross-lingual RST Discourse Parsing
2017
Discourse parsing is an integral part of understanding information flow and argumentative structure in documents. Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank. However, discourse treebanks for …
-
Cross-lingual and cross-domain discourse segmentation of entire documents
2017
Discourse segmentation is a crucial step in building end-to-end discourse parsers. However, discourse segmenters only exist for a few languages and domains. Typically they only detect intra-sentential segment boundaries, assuming gold standard sentence and token …
-
Does syntax help discourse segmentation? Not so much
2017
Discourse segmentation is the first step in building discourse parsers. Most work on discourse segmentation does not scale to real-world discourse parsing across languages, for two reasons: (i) models rely on constituent trees, and (ii) …
-
Model-based annotation of coreference
2019 · arXiv (Cornell University)
Humans do not make inferences over texts, but over models of what texts are about. When annotators are asked to annotate coreferent spans of text, it is therefore a somewhat unnatural task. This paper presents …
-
Common Sense Bias in Semantic Role Labeling
2021
Large-scale language models such as ELMo and BERT have pushed the horizon of what is possible in semantic role labeling (SRL), solving the out-of-vocabulary problem and enabling end-to-end systems, but they have also introduced significant …
-
Multi hash embeddings in spaCy
2022 · arXiv (Cornell University)
The distributed representation of symbols is one of the key technologies in machine learning systems today, playing a pivotal role in modern natural language processing. Traditional word embeddings associate a separate vector with each word. …
-
The Sensitivity of Annotator Bias to Task Definitions in Argument Mining
2022 · IT University Of Copenhagen (IT University of Copenhagen)
NLP models are dependent on the data they are trained on, including how this data is annotated. NLP research increasingly examines the social biases of models, but often in the light of their training data …
-
A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments
2017
While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set (sentence IDs) …
-
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm
2017
NLP tasks are often limited by scarcity of manually annotated data. In social media sentiment analysis and related tasks, researchers have therefore used binarized emoticons and specific hashtags as forms of distant supervision. Our paper …
-
A Survey of Cross-lingual Word Embedding Models
2019 · Journal of Artificial Intelligence Research
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …
-
Unsupervised Cross-Lingual Representation Learning
2019
In this tutorial, we provide a comprehensive survey of the exciting recent work on cutting-edge weakly-supervised and unsupervised cross-lingual word representations. After providing a brief history of supervised cross-lingual word representations, we focus on: 1) …
-
A Survey of Cross-lingual Word Embedding Models
2018 · Apollo (University of Cambridge)
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …