article Open access

Partitioning the data space before applying hashingusing clustering algorithms

  • Herald of Advanced Information Technology
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

This research presents a locality-sensitive hashingframework that enhances approximate nearest neighborsearch efficiency by integrating adaptive encoding trees and BERT-based clusterization. The proposed method optimizes data space partitioning before applying hashing, improving retrieval accuracy while reducing computational complexity. First, multimodal data, such as images and textual descriptions, are transformed into a unified semantic space using pre-trained bidirectional encoder representations from transformersembeddings. this ensures cross-modal consistency and facilitates high-dimensional similarity comparisons. Second, dimensionality reduction techniques like Uniform Manifold Approximation and Projectionor t-distributed stochastic neighbor embeddingare applied to mitigate the curse of dimensionality while preserving key relationships between data points. Third, an adaptive encoding tree locality-sensitive hashing encoding treeis constructed, dynamically segmenting the data space based on statistical distribution, thereby enabling efficient hierarchical clustering. Each data point is converted into a symbolic representation, allowing fast retrieval using structured hashing. Fourth, locality-sensitive hashingisapplied to the encoded dataset, leveraging p-stable distributions to maintain high search precision while reducing index size. The combination of encoding trees and Locality-Sensitive Hashingenables efficient candidate selection while minimizing search overhead. Experimental evaluations on the CarDD dataset, which includes car damage images and annotations, demonstrate that the proposed method outperforms state-of-the-art approximate nearest neighbor techniques in both indexing efficiency and retrieval accuracy. The results highlight its adaptability to large-scale, high-dimensional, and multimodal datasets, making it suitable for diagnostic models and real-time retrieval tasks.

Record transparency

Publication details

DOI
10.15276/hait.08.2025.2
OpenAlex
W4411880692
Document type
article
Language
EN
Source
Herald of Advanced Information Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.