conference-paper

A graph based method for Arabic document indexing

  • 2017 International Conference on Information and Digital Technologies (IDT)
Research footprint

At a glance

Citations
0
References
5
Comments
0
Paper overview

Öz

Extracting knowledge from text data and taking its full advantage has been an important way to reduce its computation and accelerate processing, especially for large amounts of data. Thus, different approaches and methodologies for modeling and representing textual data have been proposed. In this paper, a graph-based approach for automatic indexing of unstructured data from an Arabic corpus has been proposed. First, each document in the collection is represented by a graph. After the generation of document graph, term weighting is computed to estimate the relevance of a term to the document. The graph representation offers the advantage that it allows for a much more expressive document modeling than the standard bag of words approach, and consequently, it improves classification performance. Experimental results show that the graph based indexing method is a promising approach for semantic and contextual indexation, and outperforms statistical based method (TFIDF) by 12% in F-measure.

Record transparency

Publication details

DOI
10.1109/dt.2017.8012119
OpenAlex
W2747978292
Document type
conference-paper
Language
EN
Source
2017 International Conference on Information and Digital Technologies (IDT)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.