ملف الباحث
Mickel Liu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
2023 · arXiv (Cornell University)
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct …