J. Zico Kolter
9 papers in the PaperMetrix corpus
Papers by this author
-
Approximate Inference in Additive Factorial HMMs with Application to Energy Disaggregation
2018 · DSpace@MIT (Massachusetts Institute of Technology)
This paper considers additive factorial hidden Markov models, an extension to HMMs where the state factors into multiple independent chains, and the output is an additive function of all the hidden states. Although such models …
-
Realtime query completion via deep language models
2018
Search engine users nowadays heavily depend on query completion and correction to shape their queries. Typically, the completion is done by database lookup which does not understand the context and cannot generalize to prefixes not …
-
Black-box Adversarial Attacks with Bayesian Optimization
2019 · arXiv (Cornell University)
We focus on the problem of black-box adversarial attacks, where the aim is to generate adversarial examples using information limited to loss function evaluations of input-output pairs. We use Bayesian optimization~(BO) to specifically cater to …
-
Monotone operator equilibrium networks
2020 · arXiv (Cornell University)
Implicit-depth models such as Deep Equilibrium Networks have recently been shown to match or exceed the performance of traditional deep networks while being much more memory efficient. However, these models suffer from unstable convergence to …
-
Characterizing Datapoints via Second-Split Forgetting
2022 · arXiv (Cornell University)
Researchers investigating example hardness have increasingly focused on the dynamics by which neural networks learn and forget examples throughout training. Popular metrics derived from these dynamics include (i) the epoch at which examples are first …
-
Scaling Laws for Data Filtering—Data Curation Cannot be Compute Agnostic
2024
Vision-language models (VLMs) are trained for thousands of GPU hours on carefully selected subsets of massive web scrapes. For instance, the LAION public dataset retained only about 10% of the total crawled data. In recent …
-
An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
2018 · arXiv (Cornell University)
For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new …
-
Deep Equilibrium Models
2019 · arXiv (Cornell University)
We present a new approach to modeling sequential data: the deep equilibrium model (DEQ). Motivated by an observation that the hidden layers of many existing deep sequence models converge towards some fixed point, we propose …
-
Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
2023 · arXiv (Cornell University)
## Introduction Large language models (LLMs) are increasingly deployed in voice interfaces such as smartphones, smart speakers, and in-vehicle systems, which broadens the attack surface to the acoustic front end. **SWhisper (Sirens’ Whisper)** is the …