conference-paper

Accuracy of Unit Under Test Identification Using Latent Semantic Analysis and Latent Dirichlet Allocation

Research footprint

At a glance

Citations
6
References
22
Comments
0
Paper overview

Abstract

Identification of unit under test (UUT) from a test is often difficult and requires wider source code comprehension. By automating this process it would be possible to support the program comprehension and reduce software maintenance process. In this paper the Latent Semantic Analysis (LSA) and the Latent Dirichlet Allocation (LDA) were used which proved to be inaccurate in the UUT identification. The experiment was conducted on 5 popular projects where 1,093,730 similarity results were obtained. It was found out that the best topic number for the LSA model is from 7 to 10, the LDA model had big differences in this value, so it was not possible to define a stable value. The best UUT identification accuracy compared to manual testing has been obtained with the LSA model with result of 7.63% success, where documents were preprocessed using words splitting based on naming conventions and Java keywords removal. The accuracy of the LDA model was almost zero. Further 8 manual identification errors were discovered during the experiment.

Record transparency

Publication details

DOI
10.1109/informatics47936.2019.9119262
OpenAlex
W3035944277
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.