article

WEDA: Exploring Copyright Protection for Large Language Model Downstream Alignment

  • IEEE/ACM Transactions on Audio Speech and Language Processing
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
3
References
36
Comments
0
Paper overview

Öz

Large Language Models (LLMs) have shown incomparable representation and generalization capabilities, which have led to significant advancements in Natural Language Processing (NLP). Before deployment, the pre-trained LLMs often need to be tailored to specific downstream tasks for improved performance, which is commonly referred to as downstream alignment. This is a costly effort considering the needed manpower, training resources, and downstream-specific data. While much attention has been paid to protecting the copyright of the models themselves, the copyright protection of LLM alignment has been largely overlooked. In this paper, we present Watermark Embedding for Downstream Alignment (WEDA) scheme, which can provide effective copyright protection for two popular LLM alignment techniques parameter-efficient fine-tuning (PEFT) and in-context learning (ICL). For alignment through PEFT, we propose a Chain of Thought (CoT) based solution to embed watermarks into the PEFT weights. Furthermore, we extend this solution to safeguard alignment through ICL by utilizing the prefix-integrated CoT to watermark examples embedded within ICL prompts. We conduct an extensive experimental evaluation to demonstrate the effectiveness of our proposed scheme.

Record transparency

Publication details

DOI
10.1109/taslp.2024.3487419
OpenAlex
W4403863524
Document type
article
Language
EN
Source
IEEE/ACM Transactions on Audio Speech and Language Processing
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.