article وصول مفتوح

Benchmarking Large Language Models with a Unified Performance Ranking Metric

  • International Journal in Foundations of Computer Science & Technology
Research footprint

At a glance

الاستشهادات
6
المراجع
0
Comments
0
Paper overview

Abstract

The rapid advancements in Large Language Models (LLMs,) such as OpenAI’s GPT, Meta’s LLaMA, and Google’s PaLM, have revolutionized natural language processing and various AI-driven applications. Despite their transformative impact, a standardized metric to compare these models poses a significant challenge for researchers and practitioners. This paper addresses the urgent need for a comprehensive evaluation framework by proposing a novel performance ranking metric. Our metric integrates both qualitative and quantitative assessments to provide a holistic comparison of LLM capabilities. Through rigorous benchmarking, we analyze the strengths and limitations of leading LLMs, offering valuable insights into their relative performance. This study aims to facilitate informed decision-making in model selection and promote advances in developing more robust and efficient language models.

Record transparency

Publication details

DOI
10.5121/ijfcst.2024.14302
OpenAlex
W4401648846
Document type
article
Language
EN
Source
International Journal in Foundations of Computer Science & Technology
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.