Jacob Steinhardt
9 papers in the PaperMetrix corpus
Papers by this author
-
Unsupervised Risk Estimation Using Only Conditional Independence Structure
2016 · arXiv (Cornell University)
We show how to estimate a model's test error from unlabeled data, on distributions very different from the training distribution, while assuming only that certain conditional independencies are preserved between train and test. We do …
-
Semidefinite relaxations for certifying robustness to adversarial examples
2018 · arXiv (Cornell University)
Despite their impressive performance on diverse tasks, neural networks fail catastrophically in the presence of adversarial inputs---imperceptibly but adversarially perturbed versions of natural inputs. We have witnessed an arms race between defenders who attempt to …
-
What Would Jiminy Cricket Do? Towards Agents That Behave Morally
2021 · arXiv (Cornell University)
When making everyday decisions, people are guided by their conscience, an internal sense of right and wrong. By contrast, artificial agents are currently not endowed with a moral sense. As a consequence, they may learn …
-
Capturing Failures of Large Language Models via Human Cognitive Biases
2022 · arXiv (Cornell University)
Large language models generate complex, open-ended outputs: instead of outputting a class label they write summaries, generate dialogue, or produce working code. In order to asses the reliability of these open-ended generation systems, we aim …
-
Progress measures for grokking via mechanistic interpretability
2023 · arXiv (Cornell University)
Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that …
-
Adversaries Can Misuse Combinations of Safe Models
2024 · arXiv (Cornell University)
Developers try to evaluate whether an AI system can be misused by adversaries before releasing it; for example, they might test whether a model enables cyberoffense, user manipulation, or bioterrorism. In this work, we show …
-
Measuring Massive Multitask Language Understanding
2020 · arXiv (Cornell University)
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …
-
Measuring Massive Multitask Language Understanding
2021 · International Conference on Learning Representations
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …
-
Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level
2021
Larger language models have higher accuracy on average, but are they better on every single instance (datapoint)? Some work suggests larger models have higher out-ofdistribution robustness, while other work suggests they have lower accuracy on …