ملف الباحث

Alexander Toshev

4 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Multimodal Autoregressive Pre-training of Large Vision Encoders

    2024 · arXiv (Cornell University)

    We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this …

  2. Show and tell: A neural image caption generator

    2015

    Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent …

  3. Generation and Comprehension of Unambiguous Object Descriptions

    2016

    We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to …

  4. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    2022 · arXiv (Cornell University)

    Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a …