conference-paper

Extracting Tourist Attraction Entities from Text using Conditional Random Fields

Research footprint

At a glance

Citations
2
References
28
Comments
0
Paper overview

Öz

With the abundance of data available on the internet today, extracting information from texts has become critical. As tourism becomes more popular, individuals need to search for tourism-related information. To deal with this issue, named entity recognition (NER) can be applied to assist people in getting more relevant information and extracting entities from articles on the websites. This paper aims to identify entities associated with tourist destinations, such as natural, heritage, and purpose tourist attractions. We used the open dataset containing 92 articles in English. Conditional Random Fields (CRF) with various features were proposed to extract tourist attractions. Four scenarios with 13 different features were conducted to find the best NER model. Among the four scenarios, scenario 4 performed best with 97.9 precision, 94.65 recall, and 95.75 F1 scores. We found four features contribute the most to identifying the correct tag: lowercase, n-gram, previous tag, and next tag features. Based on the experiments, adding previous and next tag information could improve model performance.

Record transparency

Publication details

DOI
10.1109/icitda55840.2022.9971310
OpenAlex
W4311304580
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.