Xiaochun Cao
8 papers in the PaperMetrix corpus
Papers by this author
-
Fast Stochastic Ordinal Embedding With Variance Reduction and Adaptive Step Size
2019 · IEEE Transactions on Knowledge and Data Engineering
Learning representation from relative similarity comparisons, often called ordinal embedding, gains rising attention in recent years. Most of the existing methods are based on semi-definite programming (SDP), which is generally time-consuming and degrades the scalability, …
-
Weighted Focus-Attention Deep Network for Fine-grained Image Classification
2019
Fine-Grained Visual Classification (FGVC) is a challenging task, due to the small variation of visual representations from different categories. An effective solution is utilizing the bounding boxes centering the object parts to extract the discriminative …
-
Geometry Interaction Knowledge Graph Embeddings
2022 · arXiv (Cornell University)
Knowledge graph (KG) embeddings have shown great power in learning representations of entities and relations for link prediction tasks. Previous work usually embeds KGs into a single geometric space such as Euclidean space (zero curved), …
-
A Unified Perspective for Loss-Oriented Imbalanced Learning via Localization
2023 · arXiv (Cornell University)
Due to the inherent imbalance in real-world datasets, naïve Empirical Risk Minimization (ERM) tends to bias the learning process towards the majority classes, hindering generalization to minority classes. To rebalance the learning process, one straightforward …
-
Logit Standardization in Knowledge Distillation
2024 · arXiv (Cornell University)
Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between …
-
Promises and perils of using Transformer-based models for SE research
2024 · Neural Networks
Many Transformer-based pre-trained models for code have been developed and applied to code-related tasks. In this paper, we analyze 519 papers published on this topic during 2017-2023, examine the suitability of model architectures for different …
-
Graph Convolutional Mixture-of-Experts Learner Network for Long-Tailed Domain Generalization
2025 · IEEE Transactions on Circuits and Systems for Video Technology
The goal of single domain generalization is to use data from a single domain (source domain) to train a model, which is then deployed over several unknown domains for testing (target domains). This study introduces …
-
Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging
2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence
Multi-task learning (MTL) leverages a shared model to accomplish multiple tasks and facilitate knowledge transfer. Recent research on task arithmetic-based MTL demonstrates that merging the parameters of independently fine-tuned models can effectively achieve MTL. However, …