Researcher profile
Erik Jones
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Capturing Failures of Large Language Models via Human Cognitive Biases
2022 · arXiv (Cornell University)
Large language models generate complex, open-ended outputs: instead of outputting a class label they write summaries, generate dialogue, or produce working code. In order to asses the reliability of these open-ended generation systems, we aim …
-
Adversaries Can Misuse Combinations of Safe Models
2024 · arXiv (Cornell University)
Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulation, or bioterrorism. In this work, we show …