ملف الباحث

Harethah Abu Shairah

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. An Embarrassingly Simple Defense Against LLM Abliteration Attacks

    2025 · arXiv (Cornell University)

    Large language models (LLMs) are typically aligned to refuse harmful instructions through safety fine-tuning. A recent attack, termed abliteration, identifies and suppresses the single latent direction most responsible for refusal behavior, thereby enabling models to …