preprint

Comparative Evaluation of Large Reasoning Models and Large Language Models in Radiation Oncology: Performance in Nasopharyngeal Carcinoma-Related Queries (Preprint)

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Öz

<sec> <title>BACKGROUND</title> Large language models (LLMs), such as ChatGPT-4, have demonstrated potential in supporting medical decision-making but remain limited in scenarios requiring subjective judgment and complex clinical reasoning. Emerging large reasoning models (LRMs), including DeepSeek-R1 and Grok 3, introduce enhanced reasoning capabilities that simulate human cognitive processes, yet their clinical utility remains uncertain. Nasopharyngeal carcinoma (NPC), which demands a balance between treatment efficacy and quality of life through multidisciplinary coordination, serves as an ideal context for evaluating AI-assisted decision-making. Despite growing interest in LLMs within medicine, systematic comparisons between LLMs and LRMs in addressing open-ended clinical questions in radiation oncology remain scarce. </sec> <sec> <title>OBJECTIVE</title> This study was aimed at comparatively evaluating the performance of large language models (LLMs) and large reasoning models (LRMs) in addressing clinical management problems related to nasopharyngeal carcinoma (NPC), a complex challenge in radiation oncology. </sec> <sec> <title>METHODS</title> Five models, including three LLMs (GPT-4, GPT-4o, and Gemini 2.0 Flash) and two LRMs [Deepseek-R1 and Grok 3 (Think)] — were evaluated using a novel set of 50 open-ended questions across five modules on NPC management. Using a standardized rubric, responses were independently scored in a single-blind manner by two radiation oncologists, with statistical analyses performed to assess performance differences among models. </sec> <sec> <title>RESULTS</title> The two LRMs achieved higher mean scores (range, 16.66–17.44) than the three LLMs (range, 14.04–15.54). Statistical analysis of overall performance scores revealed that Grok 3 (Think) and Deepseek-R1 significantly outperformed both ChatGPT-4 and Gemini 2.0 Flash, and ChatGPT-4o performed better than ChatGPT-4 (P = 0.047). Module-specific analysis showed that Grok 3 (Think) and Deepseek-R1 consistently received higher scores across all modules, showing particularly superior performance in the more complex areas of Multidisciplinary Treatment and Radiotherapy. In multidimensional performance evaluation, Grok 3 (Think) excelled in accuracy (84.0%) and relevance (91.6%) and Deepseek-R1 excelled in comprehensiveness (83.2%). However, all models showed varying degrees of limitations, including reliance on outdated information, susceptibility to hallucinations, and a lack of accurate source attribution. </sec> <sec> <title>CONCLUSIONS</title> LRMs outperform LLMs in addressing open-ended NPC-related clinical management questions and show advanced potential for clinical decision support in radiation oncology . However, rigorous scrutiny of each response’s credibility remains imperative. </sec>

Record transparency

Publication details

DOI
10.2196/preprints.79516
OpenAlex
W4411645895
Document type
preprint
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.