conference-paper Open access

SYN2015: Representative Corpus of Contemporary Written Czech

Research footprint

At a glance

Citations
30
References
0
Comments
0
Paper overview

Öz

The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech.SYN2015 is a sequel of the representative corpora of the SYN series that can be described as traditional (as opposed to the web-crawled corpora), featuring cleared copyright issues, well-defined composition, reliability of annotation and high-quality text processing.At the same time, SYN2015 is designed as a reflection of the variety of written Czech text production with necessary methodological and technological enhancements that include a detailed bibliographic annotation and text classification based on an updated scheme.The corpus has been produced using a completely rebuilt text processing toolchain called SynKorp.SYN2015 is lemmatized, morphologically and syntactically annotated with state-of-the-art tools.It has been published within the framework of the Czech National Corpus and it is available via the standard corpus query interface KonText at http://kontext.korpus.czas well as a dataset in shuffled format.

Record transparency

Publication details

DOI
10.63317/39we3j4t4hi9
OpenAlex
W2578181650
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.