article

Automatic Extraction System of Webpage Information

  • Journal of Changchun University of Science and Technology
  • Changchun University of Science and Technology
Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Abstract

With the rapid development of the internet age,users have put forward more requirements for search engines,content of webpage and large data processing etc. Selecting the required information from the internet information with mass data has become a new hotspot. In this paper,extensible webcrawler project- Heritrix,which is an open source and developed by Java,is extended to capture user webpage. The information collection technology is further studied. Extendibility of Heritrix is used to realize a user's capture. Through the analysis of the working process of Heritrix,module allocation and source code design,based on webpage extraction facing product information with Heritrix extendibility and webpage content analysis with Html Parser,key product information is extracted effectively,which is stored in the database for retrieval.

Record transparency

Publication details

OpenAlex
W2367255031
Document type
article
Language
EN
Source
Journal of Changchun University of Science and Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.