Chaowei Xiao
4 papers in the PaperMetrix corpus
Papers by this author
-
Generating Adversarial Examples with Adversarial Networks
2018 · arXiv (Cornell University)
Deep neural networks (DNNs) have been found to be vulnerable to adversarial examples resulting from adding small-magnitude perturbations to inputs. Such adversarial examples can mislead DNNs to produce adversary-selected results. Different attack strategies have been …
-
SecretGen: Privacy Recovery on Pre-Trained Models via Distribution Discrimination
2022 · arXiv (Cornell University)
Transfer learning through the use of pre-trained models has become a growing trend for the machine learning community. Consequently, numerous pre-trained models are released online to facilitate further research. However, it raises extensive concerns on …
-
Semantic Adversarial Attacks via Diffusion Models
2023 · arXiv (Cornell University)
Traditional adversarial attacks concentrate on manipulating clean examples in the pixel space by adding adversarial perturbations. By contrast, semantic adversarial attacks focus on changing semantic attributes of clean examples, such as color, context, and features, …
-
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
2024 · arXiv (Cornell University)
In this study, we introduce RePD, an innovative attack Retrieval-based Prompt Decomposition framework designed to mitigate the risk of jailbreak attacks on large language models (LLMs). Despite rigorous pretraining and finetuning focused on ethical alignment, …