Vertical Scaling of Subdomain-Specific Small Language Models in Complex Aerospace Engineering: Synthetic QA Generation, Fine-Tuning, and LLM-as-a-Judge Evaluation
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Large Language Models (LLMs) demonstrate strong general-purpose reasoning, but their high inference cost and latency limit deployment in specialized engineering contexts. This paper investigates whether a localized, subdomain-adapted Small Language Model (SLM) can achieve competitive performance against commercial general-purpose LLMs in aerospace engineering—a domain requiring precise terminology, physics-grounded reasoning, and correct quantitative computation across highly heterogeneous subfields. We fine-tune a 3.8B-parameter Phi-4-mini-instruct model on a programmatically generated synthetic corpus spanning 30 aerospace subdomains, ranging from Aerodynamics and Hypersonic Flow to Orbital Mechanics and Space Propulsion. We construct a held-out, human-authored 150- question evaluation benchmark (5 questions per subdomain, spanning recall, reasoning, and calculation question types) and evaluate our fine-tuned SLM against two commercial baselines, gpt-4o-mini and gpt-3.5-turbo, using an automated gpt-4o LLM-as-a-Judge protocol scored across Accuracy, Completeness, and Reasoning Quality (1–5 scale). Our fine-tuned SLM achieves an overall average score of 4.11/5.00, compared to 4.46/5.00 for gpt-4o-mini and 4.19/5.00 for gpt-3.5-turbo. While the SLM does not surpass either commercial baseline in aggregate, it outright wins in 4 of 30 subdomains (Aerodynamics, Aeroelasticity, Aircraft Maintenance, Hypersonic Aerodynamics) and shows a relative strength in calculationtype questions (4.33, exceeding both baselines), while lagging most on open-ended recall questions (4.08). We report these results transparently, including where fine-tuning did not close the gap with larger commercial models, and discuss what this implies for the subdomainpurity thesis in a physics-heavy engineering domain as opposed to the text-classification-heavy regulatory domain examined in our prior work. All Q&A generation scripts, the evaluation benchmark, and the evaluation harness are available at the links below; model weights are hosted on Hugging Face.
Publication details
- DOI
- 10.5281/zenodo.21725736
- OpenAlex
- W7172012979
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.