ملف الباحث

Erik Jones

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Capturing Failures of Large Language Models via Human Cognitive Biases

    2022 · arXiv (Cornell University)

    Large language models generate complex, open-ended outputs: instead of outputting a class label they write summaries, generate dialogue, or produce working code. In order to asses the reliability of these open-ended generation systems, we aim …

  2. Adversaries Can Misuse Combinations of Safe Models

    2024 · arXiv (Cornell University)

    Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulation, or bioterrorism. In this work, we show …