Wengang Zhou
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual Representations
2020 · arXiv (Cornell University)
Contrastive learning has achieved great success in self-supervised visual representation learning, but existing approaches mostly ignored spatial information which is often crucial for visual representation. This paper presents heterogeneous contrastive learning (HCL), an effective approach …
-
Domain-Agnostic Prior for Transfer Semantic Segmentation
2022 · 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Unsupervised domain adaptation (UDA) is an important topic in the computer vision community. The key difficulty lies in defining a common property between the source and target domains so that the source-domain features can align …
-
Mastering the Game of 3v3 Snakes with Rule-Enhanced Multi-Agent Reinforcement Learning
2022
As a popular game around the world, Snakes has multiple modes with different settings. In this work, we are dedicated to the 3v3 Snakes, which is characterized by a complex mixture of competition and cooperation. …
-
Exploring Effective Mask Sampling Modeling for Neural Image Compression
2023 · arXiv (Cornell University)
Image compression aims to reduce the information redundancy in images. Most existing neural image compression methods rely on side information from hyperprior or context models to eliminate spatial redundancy, but rarely address the channel redundancy. …
-
MA2CL:Masked Attentive Contrastive Learning for Multi-Agent Reinforcement Learning
2023
Recent approaches have utilized self-supervised auxiliary tasks as representation learning to improve the performance and sample efficiency of vision-based reinforcement learning algorithms in single-agent settings. However, in multi-agent reinforcement learning (MARL), these techniques face challenges …
-
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
2024 · arXiv (Cornell University)
The advent of Large Multimodal Models (LMMs) has sparked a surge in research aimed at harnessing their remarkable reasoning abilities. However, for understanding text-rich images, challenges persist in fully leveraging the potential of LMMs, and …
-
Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
2024 · arXiv (Cornell University)
Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, …