article Open access

A semi-automatic data integration process of heterogeneous databases

  • Pattern Recognition Letters
  • Elsevier BV
Research footprint

At a glance

Citations
16
References
37
Comments
0
Paper overview

Öz

One of the most difficult issues today, is the integration of data from various sources. Thus, it arises the need of automatic Data Integration (DI) methods. However, in the literature there are fully automatic or semi-automatic DI techniques, but they require the involvement of IT-experts with specific domain skills. In this paper we present a novel DI methodology for which it is not required the involvement of IT-experts; in this methodology syntactically/semantically similar entities present in the sources are merged, by exploiting an information retrieval technique, a clustering method and a trained neural network. Although the suggested process is completely automated, we planned some interactions with the Company Manager, a figure who is not required to have IT-skills, but whose only contribution will be to define limits and tolerance thresholds during the DI process, based on the interests of the company. The validity of the proposed approach showed an integration accuracy between 99%−100%.

Record transparency

Publication details

DOI
10.1016/j.patrec.2023.01.007
OpenAlex
W4315929032
Document type
article
Language
EN
Source
Pattern Recognition Letters
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.