Extracting Tourist Attraction Entities from Text using Conditional Random Fields
At a glance
- Citations
- 2
- References
- 28
- Comments
- 0
Öz
With the abundance of data available on the internet today, extracting information from texts has become critical. As tourism becomes more popular, individuals need to search for tourism-related information. To deal with this issue, named entity recognition (NER) can be applied to assist people in getting more relevant information and extracting entities from articles on the websites. This paper aims to identify entities associated with tourist destinations, such as natural, heritage, and purpose tourist attractions. We used the open dataset containing 92 articles in English. Conditional Random Fields (CRF) with various features were proposed to extract tourist attractions. Four scenarios with 13 different features were conducted to find the best NER model. Among the four scenarios, scenario 4 performed best with 97.9 precision, 94.65 recall, and 95.75 F1 scores. We found four features contribute the most to identifying the correct tag: lowercase, n-gram, previous tag, and next tag features. Based on the experiments, adding previous and next tag information could improve model performance.
Publication details
- DOI
- 10.1109/icitda55840.2022.9971310
- OpenAlex
- W4311304580
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.