conference-paper
A Serbian Question Answering Dataset Created by Using the Web Scraping Technique
Research footprint
At a glance
- Citations
- 0
- References
- 17
- Comments
- 0
Paper overview
Abstract
Every artificial intelligence task requires a particular dataset to train the model and test it. As the expansion of the field of AI accelerates, data is becoming a critical resource. Natural language processing is a specific field in artificial intelligence that requires separate datasets for each task and each processed language. This paper describes the process of collecting a dataset for a question answering system in the Serbian language. Data collection was achieved using the Web scraping method. The Web scraper was implemented in the Python programming language. The resulting dataset contains 16374 questions and answers in 6 different fields: history, biology, geography, physics, chemistry, and mathematics.
Record transparency
Publication details
- DOI
- 10.1109/icest58410.2023.10187370
- OpenAlex
- W4385235766
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.