Janardhan Kulkarni
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Privately Aligning Language Models with Reinforcement Learning
2023 · arXiv (Cornell University)
Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGPT. In this work, we initiate the study …
-
TinyGSM: achieving >80% on GSM8k with small language models
2023 · arXiv (Cornell University)
Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving grade school math, the smallest model size so far required to …
-
Differentially Private Training of Mixture of Experts Models
2024 · arXiv (Cornell University)
This position paper investigates the integration of Differential Privacy (DP) in the training of Mixture of Experts (MoE) models within the field of natural language processing. As Large Language Models (LLMs) scale to billions of …
-
Differentially Private Fine-tuning of Language Models
2024 · Journal of Privacy and Confidentiality
We give simpler, sparser, and faster algorithms for differentially private fine-tuning of large-scale pre-trained language models, which achieve the state-of-the-art privacy versus utility tradeoffs on many standard NLP tasks. We propose a meta-framework for this …