ملف الباحث

Florian Tramèr

4 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy

    2022 · arXiv (Cornell University)

    Studying data memorization in neural language models helps us understand the risks (e.g., to privacy or copyright) associated with models regurgitating training data and aids in the development of countermeasures. Many prior works -- and …

  2. Backdoor Attacks for In-Context Learning with Language Models

    2023 · arXiv (Cornell University)

    Because state-of-the-art language models are expensive to train, most practitioners must make use of one of the few publicly available language models or language model APIs. This consolidation of trust increases the potency of backdoor …

  3. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

    2025 · arXiv (Cornell University)

    How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful knowledge or remotely triggering malicious actions, respectively) are typically …

  4. Quantifying Memorization Across Neural Language Models

    2022 · arXiv (Cornell University)

    Large language models (LMs) have been shown to memorize parts of their training data, and when prompted appropriately, they will emit the memorized training data verbatim. This is undesirable because memorization violates privacy (exposing user …