Digital Library XML Metadata Storage based on Data Mining Algorithm
At a glance
- Citations
- 0
- References
- 8
- Comments
- 0
Abstract
The efficient storage and retrieval of XML metadata in digital libraries face challenges such as space waste, redundancy, and low query performance under massive data volumes. Traditional XML storage methods struggle to balance scalability and efficiency. This study proposes an Apriori algorithm-based optimization framework. XML metadata undergoes preprocessing, followed by frequent itemset mining and association rule generation using Apriori. These rules guide metadata grouping and joint indexing to reduce redundancy and enhance retrieval. Experiments on datasets (10–100 GB) demonstrate significant improvements: optimized storage reduces retrieval times by 42–50% compared to traditional methods. While storage space utilization slightly decreases, the method enhances scalability and flexibility, maintaining stable performance as data grows. The approach effectively addresses storage inefficiencies and accelerates queries in large-scale environments. Limitations include computational overhead for extensive frequent itemsets. Future work will explore integration with deep learning and broader metadata formats.
Publication details
- DOI
- 10.1109/icdcece65353.2025.11034852
- OpenAlex
- W4411408883
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.