Researcher profile

Anna Korhonen

15 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Automatic Selection of Context Configurations for Improved Class-Specific Word Representations

    2017

    This paper is concerned with identifying contexts useful for training word representation models for different word classes such as adjectives (A), verbs (V), and nouns (N). We introduce a simple yet effective framework for an …

  2. Learning to Understand Phrases by Embedding the Dictionary

    2015 · arXiv (Cornell University)

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  3. Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity

    2020 · Computational Linguistics

    We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each …

  4. Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation

    2022 · arXiv (Cornell University)

    Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD - both for modular …

  5. Probing Pretrained Language Models for Lexical Semantics

    2020 · arXiv (Cornell University)

    The success of large pretrained language models (LMs) such as BERT and RoBERTa has sparked interest in probing their representations, in order to unveil what types of knowledge they implicitly capture. While prior research focused …

  6. Are Large Language Model Temporally Grounded?

    2024

    Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo Ponti, Shay Cohen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). …

  7. A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI

    2025

    Disagreement in human labeling is ubiquitous, and can be captured in human judgment distributions (HJDs).Recent research has shown that explanations provide valuable information for understanding human label variation (HLV) and large language models (LLMs) can …

  8. Learning to Understand Phrases by Embedding the Dictionary

    2016 · Transactions of the Association for Computational Linguistics

    Distributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the …

  9. SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

    2015 · Computational Linguistics

    We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies …

  10. Learning Distributed Representations of Sentences from Unlabelled Data

    2016

    Unsupervised methods for learning distributed representations of words are ubiquitous in today's NLP research, but far less is known about the best ways to learn distributed phrase or sentence representations from unlabelled data. This paper …

  11. On the Role of Seed Lexicons in Learning Bilingual Word Embeddings

    2016

    A shared bilingual word embedding space (SBWES) is an indispensable resource in a variety of cross-language NLP and IR tasks. A common approach to the SB-WES induction is to learn a mapping function between monolingual …

  12. How to Train good Word Embeddings for Biomedical NLP

    2016

    The quality of word embeddings depends on the input corpora, model architectures, and hyper-parameter settings. Using the state-of-the-art neural embedding tool word2vec and both intrinsic and extrinsic evaluations, we present a comprehensive study of how …

  13. Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints

    2017 · Transactions of the Association for Computational Linguistics

    We present Attract-Repel, an algorithm for improving the semantic quality of word vectors by injecting constraints extracted from lexical resources. Attract-Repel facilitates the use of constraints from mono- and cross-lingual resources, yielding semantically specialized cross-lingual …

  14. Isomorphic Transfer of Syntactic Structures in Cross-Lingual NLP

    2018

    The transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP. However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures. Frameworks such as Universal …

  15. On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling

    2018

    A key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. However, this ambition is largely hampered by the variation in structural and semantic properties, i.e. the typological …