ملف الباحث

Adam Gleave

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. On The Fragility of Learned Reward Functions

    2023 · arXiv (Cornell University)

    Reward functions are notoriously difficult to specify, especially for tasks with complex goals. Reward learning approaches attempt to infer reward functions from human feedback and preferences. Prior works on reward learning have mainly focused on …

  2. STACK: Adversarial Attacks on LLM Safeguard Pipelines

    2025 · arXiv (Cornell University)

    Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic and OpenAI guard their latest Opus 4 model and GPT-5 models using such defense pipelines, and other …