Peiran Wang
3 papers in the PaperMetrix corpus
Papers by this author
-
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
2024 · arXiv (Cornell University)
In this study, we introduce RePD, an innovative attack Retrieval-based Prompt Decomposition framework designed to mitigate the risk of jailbreak attacks on large language models (LLMs). Despite rigorous pretraining and finetuning focused on ethical alignment, …
-
Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies
2024
We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for users to …
-
SAP: Privacy-Preserving Fine-Tuning on Language Models with Split-and-Privatize Framework
2025
Pre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS). This …