Accuracy of Unit Under Test Identification Using Latent Semantic Analysis and Latent Dirichlet Allocation
At a glance
- Citations
- 6
- References
- 22
- Comments
- 0
Abstract
Identification of unit under test (UUT) from a test is often difficult and requires wider source code comprehension. By automating this process it would be possible to support the program comprehension and reduce software maintenance process. In this paper the Latent Semantic Analysis (LSA) and the Latent Dirichlet Allocation (LDA) were used which proved to be inaccurate in the UUT identification. The experiment was conducted on 5 popular projects where 1,093,730 similarity results were obtained. It was found out that the best topic number for the LSA model is from 7 to 10, the LDA model had big differences in this value, so it was not possible to define a stable value. The best UUT identification accuracy compared to manual testing has been obtained with the LSA model with result of 7.63% success, where documents were preprocessed using words splitting based on naming conventions and Java keywords removal. The accuracy of the LDA model was almost zero. Further 8 manual identification errors were discovered during the experiment.
Publication details
- DOI
- 10.1109/informatics47936.2019.9119262
- OpenAlex
- W3035944277
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.