article

Design and Implementation of a University Focused Crawler

  • Journal of Shanxi University
  • Shanxi University
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

How to obtain useful information from massive Web resource is crucial in the field of Web research.The main method to obtain domain-specific resources on Web is the strategy of focused crawler and this strategy only traverses pages related to the topic while neglects those irrelevant ones.However,the present strategy of focused crawler has deficiency in crawling efficiency and page quality.This article tries to make some improvements in the strategy from these two aspects,based on which we design and implement a university-oriented focused crawler system.The system uses a search strategy based on improved Context Graphs and a target page classifier based on Support Vector Machine(SVM)to acquire useful resources.The experimental results show that the system increases the harvest and accuracy of the crawling result by 10% and 8%respectively.

Record transparency

Publication details

OpenAlex
W2375180657
Document type
article
Language
EN
Source
Journal of Shanxi University
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.