conference-paper Open access

Estimation of Structural Similarity of XML Document Based on Frequency and Path

  • Advances in computer science research
  • Atlantis Press
Research footprint

At a glance

Citations
2
References
15
Comments
0
Paper overview

Abstract

With the continuous development of Internet and rich resources emerging on the Web, information retrieval based on XML has emerged; the similarity of documents is the basis of information retrieval. A method is proposed to compute similarity of XML documents based on path and frequency in the paper. XML document is expressed as a collection of tuple, the paths are extracted and delete the recurring in order to improve efficiency, tag is matched by WordNet; and then path similarity is computed by the fuzzy longest common subsequence and frequency; finally, the structure similarity between documents are calculated. Two experiments are done to show that the method is effective, the experiment 1 test structural similarity of 15 XML documents from 3 DTDs; the similarity computing is applied in the documents classification for real data sets in the experiment 2, and results show the accuracy may arrive at 100%.

Record transparency

Publication details

DOI
10.2991/emcs-16.2016.66
OpenAlex
W2305594663
Document type
conference-paper
Language
EN
Source
Advances in computer science research
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.