Researcher profile

Saaz, Mootez

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. On Inter-dataset Code Duplication and Data Leakage in Large Language Models

    2024 · arXiv (Cornell University)

    Motivation. Large language models (LLMs) have exhibited remarkable proficiency in diverse software engineering (SE) tasks. Handling such tasks typically involves acquiring foundational coding knowledge on large, general-purpose datasets during a pre-training phase, and subsequently refining …