Researcher profile
Daniel M. Ziegler
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Fine-Tuning Language Models from Human Preferences
2019 · arXiv (Cornell University)
Reward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Most work on reward learning has used simulated …
-
Language Models are Few-Shot Learners
2020 · arXiv (Cornell University)
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still …