ملف الباحث
Jean-Baptiste Gaya
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Building a Subspace of Policies for Scalable Continual Learning
2022 · arXiv (Cornell University)
The ability to continuously acquire new knowledge and skills is crucial for autonomous agents. Existing methods are typically based on either fixed-size models that struggle to learn a large number of diverse behaviors, or growing-size …
-
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
2023 · arXiv (Cornell University)
Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage. Yet the imperfections in …