conference-paper وصول مفتوح

Do LLMs Learn Structure or Names? A Study on the Robustness of Design-Pattern Detection to Identifier Shifts

Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

Large Language Models (LLMs) have recently achieved strong results for automated object-oriented Design-Pattern Detection (DPD). However, it remains unclear whether these gains come from learning pattern-defining structure (e.g., inheritance, delegation, object creation) or from exploiting identifier names (class, method, variable names) and pattern annotations. This distinction matters in practice as names are often incomplete, misleading, project-specific, or obfuscated. Developers may even label code with a given pattern name despite incorrect implementations. In such cases, an LLM that relies on names can become confidently wrong. Conversely, if code is already clean and explicity annotated with the correct pattern name, the detection model may not really be needed. In both scenarios, reliance on identifier names may defeat the purpose of design pattern detection. We conduct a study based on five LLM encoders and a dataset containing 1300 design-pattern instances collected from 214 Java projects to determine whether current LLM-based DPD genuinely learns structure or rely on name-based shortcuts. Our results show that models trained on raw code are dominated by identifier-driven attributions and suffer substantial performance degradation when tested on code with different identifier names. We also find that code anonymisation replacing identifier names by randomly set names increases reliance on code structure but does not fully eliminate identifier dependence, leaving a residual robustness gap under out-of-sample naming conventions. This study highlights the importance of understanding the features learned by LLMs before adopting them in practice.

Record transparency

Publication details

DOI
10.1145/3803846.3807465
OpenAlex
W7154834473
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.