Researcher profile

Osman, Mohamed

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. PAM: Training Policy-Aligned Moderation Filters at Scale

    2025 · arXiv (Cornell University)

    Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world …