Researcher profile
Wang, Wei
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
AJF: Adaptive Jailbreak Framework Based on the Comprehension Ability of Black-Box Large Language Models
2025 · ArXiv.org
Recent advancements in adversarial jailbreak attacks have exposed critical vulnerabilities in Large Language Models (LLMs), enabling the circumvention of alignment safeguards through increasingly sophisticated prompt manipulations. Our experiments find that the effectiveness of jailbreak strategies …