Towards Automatic Topical Classification of LOD Datasets
At a glance
- Citations
- 10
- References
- 18
- Comments
- 0
Abstract
The datasets that are part of the Linking Open Data cloud \ndiagramm (LOD cloud) are classified into the following topical \ncategories: media, government, publications, life sciences, \ngeographic, social networking, user-generated content, \nand cross-domain. The topical categories were manually \nassigned to the datasets. In this paper, we investigate to \nwhich extent the topical classification of new LOD datasets \ncan be automated using machine learning techniques and the \nexisting annotations as supervision. We conducted experiments \nwith different classification techniques and different \nfeature sets. The best classification technique/feature set \ncombination reaches an accuracy of 81.62% on the task of \nassigning one out of the eight classes to a given LOD dataset. \nA deeper inspection of the classification errors reveals problems \nwith the manual classification of datasets in the current \nLOD cloud. \n
Publication details
- OpenAlex
- W2265073338
- Document type
- conference-paper
- Language
- EN
- Source
- BOA (University of Milano-Bicocca)
- Last metadata update
Comments
Log in to join the discussion.