Shuicheng Yan
6 papers in the PaperMetrix corpus
Papers by this author
-
Mugs: A Multi-Granular Self-Supervised Learning Framework
2022 · arXiv (Cornell University)
In self-supervised learning, multi-granular features are heavily desired though rarely investigated, as different downstream tasks (e.g., general and fine-grained classification) often require different or multi-granular features, e.g.~fine- or coarse-grained one or their mixture. In this …
-
Bag of Tricks for Training Data Extraction from Language Models
2023 · arXiv (Cornell University)
With the advance of language models, privacy protection is receiving more attention. Training data extraction is therefore of great importance, as it can serve as a potential tool to assess privacy leakage. However, due to …
-
Group $K$-Means
2015 · arXiv (Cornell University)
We study how to learn multiple dictionaries from a dataset, and approximate any data point by the sum of the codewords each chosen from the corresponding dictionary. Although theoretically low approximation errors can be achieved …
-
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
2024 · arXiv (Cornell University)
Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though promising, models trained via contrastive learning on text-image pairs often neglect mid/low-level visual cues and …
-
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
2025 · arXiv (Cornell University)
Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, …
-
ConvBERT: Improving BERT with Span-based Dynamic Convolution
2020 · arXiv (Cornell University)
Pre-trained language models like BERT and its variants have recently achieved impressive performance in various natural language understanding tasks. However, BERT heavily relies on the global self-attention block and thus suffers large memory footprint and …