ملف الباحث

Huseyin A. Inan

5 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Training Data Leakage Analysis in Language Models

    2021 · arXiv (Cornell University)

    Recent advances in neural network based language models lead to successful deployments of such models, improving user experience in various applications. It has been demonstrated that strong performance of language models comes along with the …

  2. Privately Aligning Language Models with Reinforcement Learning

    2023 · arXiv (Cornell University)

    Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGPT. In this work, we initiate the study …

  3. Differentially Private Training of Mixture of Experts Models

    2024 · arXiv (Cornell University)

    This position paper investigates the integration of Differential Privacy (DP) in the training of Mixture of Experts (MoE) models within the field of natural language processing. As Large Language Models (LLMs) scale to billions of …

  4. Differentially Private Fine-tuning of Language Models

    2024 · Journal of Privacy and Confidentiality

    We give simpler, sparser, and faster algorithms for differentially private fine-tuning of large-scale pre-trained language models, which achieve the state-of-the-art privacy versus utility tradeoffs on many standard NLP tasks. We propose a meta-framework for this …

  5. Controllable Synthetic Clinical Note Generation with Privacy Guarantees

    2024 · arXiv (Cornell University)

    In the field of machine learning, domain-specific annotated data is an invaluable resource for training effective models. However, in the medical domain, this data often includes Personal Health Information (PHI), raising significant privacy concerns. The …