Select-Them-Repair-Them: Automatically Optimize Pseudocode-to-Code Data
At a glance
- Citations
- 0
- References
- 29
- Comments
- 0
Abstract
This paper considers the pseudocode-to-code generation task. Training the pseudocode-to-code generation model requires a set of labeled (pseudocode, code) pairs. However, manually annotated data pairs are costly. It is also timeconsuming to check and correct annotation errors in the annotation process. This paper proposes a novel Select-ThemRepair-Them (STRT) approach to select and repair pseudocode annotation errors to optimize data, which is able to explore pseudocode-to-code generation tasks better. This work has two key ideas: (1) although it is difficult to judge whether the pseudocode is correct, a critic is able to judge the correctness of the code corresponding to the pseudocode. (2) The pseudocode-to code generation model and code-to-pseudocode generation model are constructed to filter and correct data pairs. This paper applies the STRT approach to the SPoC dataset to obtain a new SPoC2022. The experimental results of the existing pseudocode-to-code generation methods on the SPoC2022 have increased by about 10% to 30%, which shows that the STRT is powerful.
Publication details
- DOI
- 10.1109/icbda57405.2023.10104712
- OpenAlex
- W4366725158
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.