conference-paper Open access

Jambu: A historical linguistic database for South Asian languages

Research footprint

At a glance

Citations
0
References
50
Comments
0
Paper overview

Abstract

We introduce JAMBU, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes nearly 287k lemmata from 602 lects, grouped together in 23k sets of cognates. We outline the data wrangling necessary to compile the dataset and train neural models for reflex prediction on the Indo- Aryan subset of the data. We hope that JAMBU is an invaluable resource for all historical linguists and Indologists, and look towards further improvement and expansion of the database.

Record transparency

Publication details

DOI
10.18653/v1/2023.sigmorphon-1.8
OpenAlex
W4385718243
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.