article

COSMIC: A Novel Contextualized Orientation Similarity Metric Incorporating Consistency for NLG Assessment

  • IEEE Transactions on Artificial Intelligence
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

الاستشهادات
0
المراجع
20
Comments
0
Paper overview

Abstract

The field of Natural Language Generation (NLG) has undergone remarkable expansion, largely enabled by enhanced model architectures, affordable computing, and availability of large datasets. With NLG systems finding increasing adoption across many applications, the imperative to evaluate their performance has grown exponentially. However, relying solely on human evaluation for evaluation is non-scalable. To address this challenge, it is important to explore more scalable evaluation methodologies that can ensure the continued development and efficacy of NLG systems. Presently, only a few automated evaluation metrics are commonly utilized, with BLEU and ROUGE being the predominant choices. Yet, these metrics have faced criticism for their limited correlation with human judgment, their focus on surface-level similarity, and their tendency to overlook semantic nuances. While transformer metrics have been introduced to capture semantic similarity, our study reveals scenarios where even these metrics fail. Considering these limitations, we propose and validate a novel metric called “COSMIC”, which incorporates contradiction detection with contextual embedding similarity. To illustrate these limitations and showcase the performance of COSMIC, we conducted a case study using a fine-tuned LLAMA model to transform questions and short answers to declarative sentences. This task, despite its significance in generating Natural Language Inference datasets, has not received widespread exploration since 2018. Results show that COSMIC can capture cases of contradiction between the reference and generated text, while staying highly correlated with embeddings similarity when the reference and generated text are consistent and semantically similar. BLEU, ROUGE, and most transformer-based metrics demonstrate an inability to identify contradictions.

Record transparency

Publication details

DOI
10.1109/tai.2025.3574292
OpenAlex
W4410771892
Document type
article
Language
EN
Source
IEEE Transactions on Artificial Intelligence
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.