conference-paper

DOM-based keyword extraction from web pages

Research footprint

At a glance

Citations
7
References
31
Comments
0
Paper overview

Abstract

We present D-rank, an unsupervised, language and domain independent method for automatically extracting keywords from a single web page. The method does not use any corpus, and relies only on the information and features on the web page including page URL, word frequency, title, hyperlinks, and headers, which are extracted from DOM tree of the page. Different scores are assigned to the words according to their importance that is specified by their positions in the web page. Experimental results on web pages in three different languages show the effectiveness of the proposed method.

Record transparency

Publication details

DOI
10.1145/3371425.3371495
OpenAlex
W2995983043
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.