Researcher profile

Weng, Jiaqi

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models

    2025 · arXiv (Cornell University)

    Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: generalization and accuracy. While learning-based methods trained …