Synthetic Data Generation for Supervised Fine-Tuning: A Comprehensive Survey
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
This survey provides a comprehensive analysis of Synthetic Data Generation (SDG) for Supervised Fine-Tuning (SFT), a transformative shift in the Large Language Model (LLM) alignment lifecycle. As the demand for high-quality instructional data outpaces human annotation capacity, synthetic data has emerged as the primary vehicle for scaling model capabilities across natural conversation, task-oriented dialogue, tool-calling, and agentic workflows. We introduce a multi-dimensional taxonomy of synthetic SFT data and detail the end-to-end pipeline architecture, from schema design to quality control. We survey a spectrum of generation methods—ranging from deterministic template-based approaches to sophisticated multi-agent agentic pipelines—and provide actionable implementation recipes for 15 industry verticals, including healthcare, banking, and telecom. Furthermore, we examine the critical roles of consistent user simulation, automated verification, and multi-layered evaluation in ensuring dataset integrity. Finally, we address systemic risks such as model collapse, bias amplification, and regulatory compliance, offering a governance framework for the responsible deployment of synthetic pipelines. This work serves as a foundational reference for practitioners and researchers navigating the transition to synthetic-centric model alignment.
Publication details
- DOI
- 10.5281/zenodo.19673643
- OpenAlex
- W7155031460
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.