Ivan Vulić
16 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Automatic Selection of Context Configurations for Improved Class-Specific Word Representations
2017
This paper is concerned with identifying contexts useful for training word representation models for different word classes such as adjectives (A), verbs (V), and nouns (N). We introduce a simple yet effective framework for an …
-
Failure points in the PKI architecture
2017 · Vojnotehnicki glasnik
Over the last 20 years, the PKI architecture has found its vast application, especially in the fields which require the establishment of a security infrastructure. Given that the purpose of this architecture is to be …
-
Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity
2020 · Computational Linguistics
We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each …
-
Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation
2022 · arXiv (Cornell University)
Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD - both for modular …
-
Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual Retrieval
2022 · arXiv (Cornell University)
State-of-the-art neural (re)rankers are notoriously data-hungry which -- given the lack of large-scale training data in languages other than English -- makes them rarely used in multilingual and cross-lingual retrieval settings. Current approaches therefore commonly …
-
Monolingual and Cross-Lingual Information Retrieval Models Based on (Bilingual) Word Embeddings
2015
We propose a new unified framework for monolingual (MoIR) and cross-lingual information retrieval (CLIR) which relies on the induction of dense real-valued word vectors known as word embeddings (WE) from comparable data. To this end, …
-
On the Role of Seed Lexicons in Learning Bilingual Word Embeddings
2016
A shared bilingual word embedding space (SBWES) is an indispensable resource in a variety of cross-language NLP and IR tasks. A common approach to the SB-WES induction is to learn a mapping function between monolingual …
-
Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints
2017 · Transactions of the Association for Computational Linguistics
We present Attract-Repel, an algorithm for improving the semantic quality of word vectors by injecting constraints extracted from lexical resources. Attract-Repel facilitates the use of constraints from mono- and cross-lingual resources, yielding semantically specialized cross-lingual …
-
A Survey of Cross-lingual Word Embedding Models
2019 · Journal of Artificial Intelligence Research
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …
-
Isomorphic Transfer of Syntactic Structures in Cross-Lingual NLP
2018
The transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP. However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures. Frameworks such as Universal …
-
On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling
2018
A key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. However, this ambition is largely hampered by the variation in structural and semantic properties, i.e. the typological …
-
Training Neural Response Selection for Task-Oriented Dialogue Systems
2019 · arXiv (Cornell University)
Despite their popularity in the chatbot literature, retrieval-based models have had modest impact on task-oriented dialogue systems, with the main obstacle to their application being the low-data regime of most task-oriented dialogue tasks. Inspired by …
-
JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages
2019
Viable cross-lingual transfer critically depends on the availability of parallel texts. Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing. We introduce JW300, a parallel corpus of over 300 languages with …
-
Unsupervised Cross-Lingual Representation Learning
2019
In this tutorial, we provide a comprehensive survey of the exciting recent work on cutting-edge weakly-supervised and unsupervised cross-lingual word representations. After providing a brief history of supervised cross-lingual word representations, we focus on: 1) …
-
From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers
2020
bak probing, til beskrivelser av transformerarkitekturen og debatten rundt tolkbarhet av oppmerksomhet. I hvert introduksjonskapittel prver vi derfor kontekstualisere vrt eget arbeid.
-
A Survey of Cross-lingual Word Embedding Models
2018 · Apollo (University of Cambridge)
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we …