Researcher profile

Kun Kuang

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning

    2021

    Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process …

  2. Collaborative Semantic Aggregation and Calibration for Federated Domain Generalization

    2021 · arXiv (Cornell University)

    Domain generalization (DG) aims to learn from multiple known source domains a model that can generalize well to unknown target domains. The existing DG methods usually exploit the fusion of shared multi-source data to train …

  3. Focus-aware Response Generation in Inquiry Conversation

    2023

    Inquiry conversation is a common form of conversation that aims to complete the investigation (e.g., court hearing, medical consultation and police interrogation) during which a series of focus shifts occurs. While many models have been …

  4. MEDOE: A Multi-Expert Decoder and Output Ensemble Framework for Long-tailed Semantic Segmentation

    2023 · arXiv (Cornell University)

    Long-tailed distribution of semantic categories, which has been often ignored in conventional methods, causes unsatisfactory performance in semantic segmentation on tail categories. In this paper, we focus on the problem of long-tailed semantic segmentation. Although …

  5. Transferring Causal Mechanism over Meta-representations for Target-Unknown Cross-domain Recommendation

    2024 · ACM Transactions on Information Systems

    Tackling the pervasive issue of data sparsity in recommender systems, we present an insightful investigation into the burgeoning area of non-overlapping cross-domain recommendation, a technique that facilitates the transfer of interaction knowledge across domains without …

  6. More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs

    2024 · arXiv (Cornell University)

    The performance on general tasks decreases after Large Language Models (LLMs) are fine-tuned on domain-specific tasks, the phenomenon is known as Catastrophic Forgetting (CF). However, this paper presents a further challenge for real application of …

  7. RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

    2024 · arXiv (Cornell University)

    Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in …

  8. Forward Once for All: Structural Parameterized Adaptation for Efficient Cloud-coordinated On-device Recommendation

    2025

    In cloud-centric recommender system, regular data exchanges between user devices and cloud could potentially elevate bandwidth demands and privacy risks. On-device recommendation emerges as a viable solution by performing reranking locally to alleviate these concerns. …

  9. General information metrics for improving AI model training efficiency

    2025 · Artificial Intelligence Review

    Abstract To address the growing size of AI model training data and the lack of a universal data selection methodology–factors that significantly drive up training costs–this paper presents the General Information Metrics Evaluation (GIME) method. …