ملف الباحث

Dan Hendrycks

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations

    2018 · arXiv (Cornell University)

    In this paper we establish rigorous benchmarks for image classifier robustness. Our first benchmark, ImageNet-C, standardizes and expands the corruption robustness topic, while showing which classifiers are preferable in safety-critical applications. Unlike recent robustness research, …

  2. A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges

    2021 · arXiv (Cornell University)

    Machine learning models often encounter samples that are diverged from the training distribution. Failure to recognize an out-of-distribution (OOD) sample, and consequently assign that sample to an in-class label significantly compromises the reliability of a …

  3. What Would Jiminy Cricket Do? Towards Agents That Behave Morally

    2021 · arXiv (Cornell University)

    When making everyday decisions, people are guided by their conscience, an internal sense of right and wrong. By contrast, artificial agents are currently not endowed with a moral sense. As a consequence, they may learn …

  4. EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

    2025 · arXiv (Cornell University)

    As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced reasoning and knowledge …

  5. Measuring Massive Multitask Language Understanding

    2020 · arXiv (Cornell University)

    We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …

  6. Measuring Massive Multitask Language Understanding

    2021 · International Conference on Learning Representations

    We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models …