Researcher profile
Enes Altınışık
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
PAM: Training Policy-Aligned Moderation Filters at Scale
2025 · arXiv (Cornell University)
Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world …