Dan Hendrycks
6 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
2018 · arXiv (Cornell University)
In this paper we establish rigorous benchmarks for image classifier robustness. Our first benchmark, ImageNet-C, standardizes and expands the corruption robustness topic, while showing which classifiers are preferable in safety-critical applications. Unlike recent robustness research, …
-
A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges
2021 · arXiv (Cornell University)
Machine learning models often encounter samples that are diverged from the training distribution. Failure to recognize an out-of-distribution (OOD) sample, and consequently assign that sample to an in-class label significantly compromises the reliability of a …
-
What Would Jiminy Cricket Do? Towards Agents That Behave Morally
2021 · arXiv (Cornell University)
When making everyday decisions, people are guided by their conscience, an internal sense of right and wrong. By contrast, artificial agents are currently not endowed with a moral sense. As a consequence, they may learn …
-
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
2025 · arXiv (Cornell University)
As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced reasoning and knowledge …
-
Measuring Massive Multitask Language Understanding
2020 · arXiv (Cornell University)
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …
-
Measuring Massive Multitask Language Understanding
2021 · International Conference on Learning Representations
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …