Huseyin A. Inan
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Training Data Leakage Analysis in Language Models
2021 · arXiv (Cornell University)
Recent advances in neural network based language models lead to successful deployments of such models, improving user experience in various applications. It has been demonstrated that strong performance of language models comes along with the …
-
Privately Aligning Language Models with Reinforcement Learning
2023 · arXiv (Cornell University)
Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGPT. In this work, we initiate the study …
-
Differentially Private Training of Mixture of Experts Models
2024 · arXiv (Cornell University)
This position paper investigates the integration of Differential Privacy (DP) in the training of Mixture of Experts (MoE) models within the field of natural language processing. As Large Language Models (LLMs) scale to billions of …
-
Differentially Private Fine-tuning of Language Models
2024 · Journal of Privacy and Confidentiality
We give simpler, sparser, and faster algorithms for differentially private fine-tuning of large-scale pre-trained language models, which achieve the state-of-the-art privacy versus utility tradeoffs on many standard NLP tasks. We propose a meta-framework for this …
-
Controllable Synthetic Clinical Note Generation with Privacy Guarantees
2024 · arXiv (Cornell University)
In the field of machine learning, domain-specific annotated data is an invaluable resource for training effective models. However, in the medical domain, this data often includes Personal Health Information (PHI), raising significant privacy concerns. The …