Researcher profile
Adam Gleave
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
On The Fragility of Learned Reward Functions
2023 · arXiv (Cornell University)
Reward functions are notoriously difficult to specify, especially for tasks with complex goals. Reward learning approaches attempt to infer reward functions from human feedback and preferences. Prior works on reward learning have mainly focused on …
-
STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025 · arXiv (Cornell University)
Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic and OpenAI guard their latest Opus 4 model and GPT-5 models using such defense pipelines, and other …