SYN2015: Representative Corpus of Contemporary Written Czech
At a glance
- Citations
- 30
- References
- 0
- Comments
- 0
Abstract
The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech.SYN2015 is a sequel of the representative corpora of the SYN series that can be described as traditional (as opposed to the web-crawled corpora), featuring cleared copyright issues, well-defined composition, reliability of annotation and high-quality text processing.At the same time, SYN2015 is designed as a reflection of the variety of written Czech text production with necessary methodological and technological enhancements that include a detailed bibliographic annotation and text classification based on an updated scheme.The corpus has been produced using a completely rebuilt text processing toolchain called SynKorp.SYN2015 is lemmatized, morphologically and syntactically annotated with state-of-the-art tools.It has been published within the framework of the Czech National Corpus and it is available via the standard corpus query interface KonText at http://kontext.korpus.czas well as a dataset in shuffled format.
Publication details
- DOI
- 10.63317/39we3j4t4hi9
- OpenAlex
- W2578181650
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.