ملف الباحث

Jeffrey Wu

3 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Fine-Tuning Language Models from Human Preferences

    2019 · arXiv (Cornell University)

    Reward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Most work on reward learning has used simulated …

  2. Scaling Laws for Neural Language Models

    2020 · arXiv (Cornell University)

    This paper develops a transport-validity theory for agentic AI interventions that are first screened on small systems and later considered for frontier-scale deployment. Rather than predicting absolute frontier performance, it asks when a comparative gain …

  3. Language Models are Few-Shot Learners

    2020 · arXiv (Cornell University)

    Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still …