Researcher profile
Harethah Abu Shairah
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
2025 · arXiv (Cornell University)
Large language models (LLMs) are typically aligned to refuse harmful instructions through safety fine-tuning. A recent attack, termed abliteration, identifies and suppresses the single latent direction most responsible for refusal behavior, thereby enabling models to …