Classification of the Approaches of the Near Duplicate Document Detection and Elimination
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
The area of identification and removal of near duplicate document is important to research as near duplicate pages increase overhead on web. It increases storage space and indexing cost, crawlers produces same type of results and make searching process ineffective. Identical pages on web exists because it contains huge volume of records. Many researches have been done in this area to solve the problem but the problem still exists. The researchers have studied the problem from different perspectives and tried to formulate solutions, however the problem is intensifying as new pages are added to the web. This paper studies previous research work and classify those algorithms and approaches with the intention of structuring the area of duplicate document finding.
Publication details
- OpenAlex
- W2472412549
- Document type
- article
- Language
- EN
- Source
- Research Journal of Information Technology & Software Management
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.