preprint Open access

Vertical Scaling of Subdomain-Specific Small Language Models in Complex Aerospace Engineering: Synthetic QA Generation, Fine-Tuning, and LLM-as-a-Judge Evaluation

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Large Language Models (LLMs) demonstrate strong general-purpose reasoning, but their high inference cost and latency limit deployment in specialized engineering contexts. This paper investigates whether a localized, subdomain-adapted Small Language Model (SLM) can achieve competitive performance against commercial general-purpose LLMs in aerospace engineering—a domain requiring precise terminology, physics-grounded reasoning, and correct quantitative computation across highly heterogeneous subfields. We fine-tune a 3.8B-parameter Phi-4-mini-instruct model on a programmatically generated synthetic corpus spanning 30 aerospace subdomains, ranging from Aerodynamics and Hypersonic Flow to Orbital Mechanics and Space Propulsion. We construct a held-out, human-authored 150- question evaluation benchmark (5 questions per subdomain, spanning recall, reasoning, and calculation question types) and evaluate our fine-tuned SLM against two commercial baselines, gpt-4o-mini and gpt-3.5-turbo, using an automated gpt-4o LLM-as-a-Judge protocol scored across Accuracy, Completeness, and Reasoning Quality (1–5 scale). Our fine-tuned SLM achieves an overall average score of 4.11/5.00, compared to 4.46/5.00 for gpt-4o-mini and 4.19/5.00 for gpt-3.5-turbo. While the SLM does not surpass either commercial baseline in aggregate, it outright wins in 4 of 30 subdomains (Aerodynamics, Aeroelasticity, Aircraft Maintenance, Hypersonic Aerodynamics) and shows a relative strength in calculationtype questions (4.33, exceeding both baselines), while lagging most on open-ended recall questions (4.08). We report these results transparently, including where fine-tuning did not close the gap with larger commercial models, and discuss what this implies for the subdomainpurity thesis in a physics-heavy engineering domain as opposed to the text-classification-heavy regulatory domain examined in our prior work. All Q&A generation scripts, the evaluation benchmark, and the evaluation harness are available at the links below; model weights are hosted on Hugging Face.

Record transparency

Publication details

DOI
10.5281/zenodo.21725736
OpenAlex
W7172012979
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.