Researcher profile

Guillaume Couairon

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

    2023 · arXiv (Cornell University)

    Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage. Yet the imperfections in …