article Open access

A large and evolving cognate database

  • Language Resources and Evaluation
  • Springer Science+Business Media
Research footprint

At a glance

Citations
60
References
69
Comments
0
Paper overview

Abstract

Abstract We present CogNet , a large-scale, automatically-built database of sense-tagged cognates —words of common origin and meaning across languages. CogNet is continuously evolving: its current version contains over 8 million cognate pairs over 338 languages and 35 writing systems, with new releases already in preparation. The paper presents the algorithm and input resources used for its computation, an evaluation of the result, as well as a quantitative analysis of cognate data leading to novel insights on language diversity. Furthermore, as an example on the use of large-scale cross-lingual knowledge bases for improving the quality of multilingual applications, we present a case study on the use of CogNet for bilingual lexicon induction in the framework of cross-lingual transfer learning.

Record transparency

Publication details

DOI
10.1007/s10579-021-09544-6
OpenAlex
W3167873515
Document type
article
Language
EN
Source
Language Resources and Evaluation
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.