Information Retrieval: A Comparative Study Of Textual Indexing Using An Oriented Object Database (Db4O) And The Inverted File
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
The growth in the volume of text data such as books<br> and articles in libraries for centuries has imposed to establish<br> effective mechanisms to locate them. Early techniques such as<br> abstraction, indexing and the use of classification categories have<br> marked the birth of a new field of research called "Information<br> Retrieval". Information Retrieval (IR) can be defined as the task of<br> defining models and systems whose purpose is to facilitate access to<br> a set of documents in electronic form (corpus) to allow a user to find<br> the relevant ones for him, that is to say, the contents which matches<br> with the information needs of the user.<br> Most of the models of information retrieval use a specific data<br> structure to index a corpus which is called "inverted file" or "reverse<br> index".<br> This inverted file collects information on all terms over the corpus<br> documents specifying the identifiers of documents that contain the<br> term in question, the frequency of each term in the documents of the<br> corpus, the positions of the occurrences of the word...<br> In this paper we use an oriented object database (db4o) instead of<br> the inverted file, that is to say, instead to search a term in the inverted<br> file, we will search it in the db4o database.<br> The purpose of this work is to make a comparative study to see if<br> the oriented object databases may be competing for the inverse index<br> in terms of access speed and resource consumption using a large<br> volume of data.
Publication details
- DOI
- 10.5281/zenodo.1337851
- OpenAlex
- W1668273999
- Document type
- article
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.