ملف الباحث
Enes Altınışık
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
PAM: Training Policy-Aligned Moderation Filters at Scale
2025 · arXiv (Cornell University)
Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world …