Automatic Extraction System of Webpage Information
At a glance
- Citations
- 1
- References
- 0
- Comments
- 0
Abstract
With the rapid development of the internet age,users have put forward more requirements for search engines,content of webpage and large data processing etc. Selecting the required information from the internet information with mass data has become a new hotspot. In this paper,extensible webcrawler project- Heritrix,which is an open source and developed by Java,is extended to capture user webpage. The information collection technology is further studied. Extendibility of Heritrix is used to realize a user's capture. Through the analysis of the working process of Heritrix,module allocation and source code design,based on webpage extraction facing product information with Heritrix extendibility and webpage content analysis with Html Parser,key product information is extracted effectively,which is stored in the database for retrieval.
Publication details
- OpenAlex
- W2367255031
- Document type
- article
- Language
- EN
- Source
- Journal of Changchun University of Science and Technology
- Last metadata update
Comments
Log in to join the discussion.