conference-paper
Bridging pre-trained models and downstream tasks for source code understanding
Research footprint
At a glance
- Citations
- 70
- References
- 30
- Comments
- 0
Paper overview
Abstract
With the great success of pre-trained models, the pretrain-then-finetune paradigm has been widely adopted on downstream tasks for source code understanding. However, compared to costly training a large-scale model from scratch, how to effectively adapt pre-trained models to a new task has not been fully explored. In this paper, we propose an approach to bridge pre-trained models and code-related tasks. We exploit semantic-preserving transformation to enrich downstream data diversity, and help pre-trained models learn semantic features invariant to these semantically equivalent transformations. Further, we introduce curriculum learning to organize the transformed data in an easy-to-hard manner to fine-tune existing pre-trained models.
Record transparency
Publication details
- DOI
- 10.1145/3510003.3510062
- OpenAlex
- W4200633062
- Document type
- conference-paper
- Language
- EN
- Source
- Proceedings of the 44th International Conference on Software Engineering
- Last metadata update
Comments
Log in to join the discussion.