Silvio Savarese
6 papers in the PaperMetrix corpus
Papers by this author
-
Unsupervised Transductive Domain Adaptation
2016 · arXiv (Cornell University)
Supervised learning with large scale labeled datasets and deep layered models has made a paradigm shift in diverse areas in learning and recognition. However, this approach still suffers generalization issues under the presence of a …
-
Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation
2018
Xiaoxue Zang, Ashwini Pokle, Marynel Vázquez, Kevin Chen, Juan Carlos Niebles, Alvaro Soto, Silvio Savarese. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
-
Neural Architecture Search From Fréchet Task Distance.
2021 · arXiv (Cornell University)
We formulate a Frechet-type asymmetric distance between tasks based on Fisher Information Matrices. We show how the distance between a target task and each task in a given set of baseline tasks can be used …
-
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
2025
Jierui Li, Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: …
-
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
2025 · arXiv (Cornell University)
Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified …
-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2023 · arXiv (Cornell University)
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image …