Benchmark Evaluation of a Cannabinoid Drug Interaction Chatbot
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
IMPORTANCE Large Language Models (LLMs) have enabled chatbots, such as the Cannabinoid Drug Interaction (CDI) chatbot, to leverage patterns learned from training on vast datasets to formulate a response. The CDI also demonstrates adaptability in answering broader medical and other non-cannabinoid drug interaction inquiries. OBJECTIVETo evaluate the Cannabinoid Drug Interaction (CDI) Chatbot with standard benchmarking datasets and compare to other well-known LLMs. DESIGN, SETTING, AND PARTICIPANTS The CDI implements a retrieval-augmented generation (RAG) framework that integrates the Meta-Llama-3-8B-Instruct LLM for text generation and the gte-large-en-v1.5 embedding model for prompt classification. This combination of models allows for efficient hybrid information retrieval from database and web-based search results. Additionally, the CDI chatbot accesses a specialized cannabinoid drug interaction knowledge database. INTERVENTIONLLM and embedding benchmark testing was used to evaluate the performance of the CDI chatbot. MAIN OUTCOMES AND MEASURESA benchmarking dataset, containing 280 multiple-choice cannabinoid focused questions, was developed to evaluate the CDI chatbot. The benchmarking of CDI also utilized Massive Multitask Language Understanding (MMLU) datasets to establish a standard point of reference accessed through the open source LLM evaluation framework DeepEval. RESULTS When drug database context was included, cannabinoid DDI accuracy scores were more than doubled when compared to the LLM-only performance of gemma-2-9b-it and Mistral-NeMo-12B-Instruct models. The gemma-2-9b-it with retrieval from drug database and web sources, using the gte-large-en-v1.5 embedding model, achieved 80% accuracy in MMLU Clinical Knowledge (CK) and 90% accuracy in MMLU Medical Genetics (MG) tasks. These scores are comparable to the LLM-only results of the Meta-Llama-3-70B-Instruct, which scored 79.62% on MMLU Clinical Knowledge (CK) and 90% on MMLU Medical Genetics (MG) tasks. The CDI chatbot exhibited robust retrieval capabilities for cannabinoid DDI information, with 269 out of 280 context instances scoring above 0.9 for contextual precision and 278 out of 280 context instances scoring above 0.9 for contextual relevancy. The CDI chatbot achieved > 90% accuracy in correctly answering cannabinoid drug-drug interaction (DDI) questions; thereby, outperforming other LLMs tested. CONCLUSIONS AND RELEVANCE The Cannabinoid Drug Interaction (CDI) Chatbot demonstrates promising results for providing reliable and accurate DDI information for the cannabinoid class of medications. The CDI chatbot underscores the potential of a RAG framework to advance clinical decision support (CDS) to provide accurate, context-aware drug interaction information. Further validation would be needed to assess suitability for clinical decision support deployments.
Publication details
- DOI
- 10.5281/zenodo.18509541
- OpenAlex
- W7128068710
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.