conference-paper

A Serbian Question Answering Dataset Created by Using the Web Scraping Technique

Research footprint

At a glance

Citations
0
References
17
Comments
0
Paper overview

Abstract

Every artificial intelligence task requires a particular dataset to train the model and test it. As the expansion of the field of AI accelerates, data is becoming a critical resource. Natural language processing is a specific field in artificial intelligence that requires separate datasets for each task and each processed language. This paper describes the process of collecting a dataset for a question answering system in the Serbian language. Data collection was achieved using the Web scraping method. The Web scraper was implemented in the Python programming language. The resulting dataset contains 16374 questions and answers in 6 different fields: history, biology, geography, physics, chemistry, and mathematics.

Record transparency

Publication details

DOI
10.1109/icest58410.2023.10187370
OpenAlex
W4385235766
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.