Researcher profile

Juntao Dai

2 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset

    2023 · arXiv (Cornell University)

    In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct …

  2. SafeLawBench: Towards Safe Alignment of Large Language Models

    2025 · arXiv (Cornell University)

    With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns. However, there is still a lack of definitive standards for evaluating their safety due to the subjective nature …