ملف الباحث

Tom Goldstein

13 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent

    2019 · arXiv (Cornell University)

    This paper studies how neural network architecture affects the speed of training. We introduce a simple concept called gradient confusion to help formally analyze this. When gradient confusion is high, stochastic gradients produced by different …

  2. Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks

    2020 · arXiv (Cornell University)

    Data poisoning and backdoor attacks manipulate training data in order to cause models to fail during inference. A recent survey of industry practitioners found that data poisoning is the number one concern among threats ranging …

  3. Data Augmentation for Meta-Learning

    2020 · arXiv (Cornell University)

    Conventional image classifiers are trained by randomly sampling mini-batches of images. To achieve state-of-the-art performance, practitioners use sophisticated data augmentation schemes to expand the amount of training data available for sampling. In contrast, meta-learning algorithms …

  4. Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary Perspective

    2022 · arXiv (Cornell University)

    We discuss methods for visualizing neural network decision boundaries and decision regions. We use these visualizations to investigate issues related to reproducibility and generalization in neural network training. We observe that changes in model architecture …

  5. Fishing for User Data in Large-Batch Federated Learning via Gradient Magnification

    2022 · arXiv (Cornell University)

    Federated learning (FL) has rapidly risen in popularity due to its promise of privacy and efficiency. Previous works have exposed privacy vulnerabilities in the FL pipeline by recovering user data from gradient updates. However, existing …

  6. Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with\n Recurrent Networks

    2021 · arXiv (Cornell University)

    Deep neural networks are powerful machines for visual pattern recognition,\nbut reasoning tasks that are easy for humans may still be difficult for neural\nmodels. Humans possess the ability to extrapolate reasoning strategies learned\non simple problems to …

  7. Robust Optimization as Data Augmentation for Large-scale Graphs

    2020 · arXiv (Cornell University)

    Data augmentation helps neural networks generalize better by enlarging the training set, but it remains an open question how to effectively augment graph data to enhance the performance of GNNs (Graph Neural Networks). While most …

  8. Seeing in Words: Learning to Classify through Language Bottlenecks

    2023 · arXiv (Cornell University)

    Neural networks for computer vision extract uninterpretable features despite achieving high accuracy on benchmarks. In contrast, humans can explain their predictions using succinct and intuitive descriptions. To incorporate explainability into neural networks, we train a …

  9. Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs

    2024 · arXiv (Cornell University)

    Large language models can memorize and repeat their training data, causing privacy and copyright risks. To mitigate memorization, we introduce a subtle modification to the next-token training objective that we call the goldfish loss. During …

  10. Has My System Prompt Been Used? Large Language Model Prompt Membership Inference

    2025 · arXiv (Cornell University)

    Prompt engineering has emerged as a powerful technique for optimizing large language models (LLMs) for specific applications, enabling faster prototyping and improved performance, and giving rise to the interest of the community in protecting proprietary …

  11. BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

    2025 · arXiv (Cornell University)

    Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified …

  12. Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

    2025 · arXiv (Cornell University)

    A high volume of recent ML security literature focuses on attacks against aligned large language models (LLMs). These attacks may extract private information or coerce the model into producing harmful outputs. In real-world deployments, LLMs …

  13. FreeLB: Enhanced Adversarial Training for Natural Language Understanding

    2019 · arXiv (Cornell University)

    Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that …