Alexander Toshev
4 papers in the PaperMetrix corpus
Papers by this author
-
Multimodal Autoregressive Pre-training of Large Vision Encoders
2024 · arXiv (Cornell University)
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this …
-
Show and tell: A neural image caption generator
2015
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent …
-
Generation and Comprehension of Unambiguous Object Descriptions
2016
We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to …
-
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
2022 · arXiv (Cornell University)
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a …