Researcher profile

Chiyuan Zhang

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers

    2021 · arXiv (Cornell University)

    A discriminatively trained neural net classifier can fit the training data perfectly if all information about its input other than class membership has been discarded prior to the output layer. Surprisingly, past research has discovered …

  2. Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy

    2022 · arXiv (Cornell University)

    Studying data memorization in neural language models helps us understand the risks (e.g., to privacy or copyright) associated with models regurgitating training data and aids in the development of countermeasures. Many prior works -- and …

  3. Deduplicating Training Data Makes Language Models Better

    2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.

  4. Quantifying Memorization Across Neural Language Models

    2022 · arXiv (Cornell University)

    Large language models (LMs) have been shown to memorize parts of their training data, and when prompted appropriately, they will emit the memorized training data verbatim. This is undesirable because memorization violates privacy (exposing user …