article

Classification of the Approaches of the Near Duplicate Document Detection and Elimination

  • Research Journal of Information Technology & Software Management
Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

The area of identification and removal of near duplicate document is important to research as near duplicate pages increase overhead on web. It increases storage space and indexing cost, crawlers produces same type of results and make searching process ineffective. Identical pages on web exists because it contains huge volume of records. Many researches have been done in this area to solve the problem but the problem still exists. The researchers have studied the problem from different perspectives and tried to formulate solutions, however the problem is intensifying as new pages are added to the web. This paper studies previous research work and classify those algorithms and approaches with the intention of structuring the area of duplicate document finding.

Record transparency

Publication details

OpenAlex
W2472412549
Document type
article
Language
EN
Source
Research Journal of Information Technology & Software Management
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.