Comparative Study on Test Case generation using Generative AI
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Testing is an important part of Development life cycle, resource intensive and prone to human error. Over the years, various tools have been developed, with a recent surge in those based on Large Language Models (LLMs). Examples include prompting LLMs in chat applications to generate test cases or using AI agents with project context to generate tests. This paper comparatively studies the most prominent open and closed-source models available at the time of writing. This paper covers the most famous open and closed source models as of the date of writing this paper in a comparative study. This study analyzes each model’s characteristics and performance against well known metrics like average code coverage and mutation score. While closed source models Like GPT-4o are demonstrating better performance than all other models at 35.2% coverage, some open source models such as Llama 3.1 70B is demonstrating significantly p romising p erformance with 30.6% coverage, or models like DeepSeekCoderV2 16B that have a better chance of running locally in a normal home setting with 28.2% coverage. This study highlights LLM capabilities in automated test generation.
Publication details
- DOI
- 10.1109/imsa65733.2025.11166964
- OpenAlex
- W4414463696
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.