ملف الباحث

Fei Wu

21 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Ad Recommendation for Sponsored Search Engine via Composite Long-Short Term Memory

    2016

    Search engine logs contain a large amount of users' click-through data that can be leveraged as implicit indicators of relevance. In this paper we address ad recommendation problem that finding and ranking the most relevant …

  2. Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning

    2021

    Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process …

  3. Collaborative Semantic Aggregation and Calibration for Federated Domain Generalization

    2021 · arXiv (Cornell University)

    Domain generalization (DG) aims to learn from multiple known source domains a model that can generalize well to unknown target domains. The existing DG methods usually exploit the fusion of shared multi-source data to train …

  4. A General Framework for Defending Against Backdoor Attacks via Influence Graph

    2021 · arXiv (Cornell University)

    In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attacking pattern, and therefore, poisoned training examples …

  5. Combined Forecasting of Ship Heave Motion Based on Induced Ordered Weighted Averaging Operator

    2022 · IEEJ Transactions on Electrical and Electronic Engineering

    Abstract Heave motion of ships is a complex nonlinear dynamic process and cannot be accurately forecasted using a single prediction model. In this paper, an effective combined forecasting method is proposed to perform ship's heave …

  6. GPT-NER: Named Entity Recognition via Large Language Models

    2023 · arXiv (Cornell University)

    Despite the fact that large-scale Language Models (LLM) have achieved SOTA performances on a variety of NLP tasks, its performance on NER is still significantly below supervised baselines. This is due to the gap between …

  7. Pushing the Limits of ChatGPT on NLP Tasks

    2023 · arXiv (Cornell University)

    Despite the success of ChatGPT, its performances on most NLP tasks are still well below the supervised baselines. In this work, we looked into the causes, and discovered that its subpar performance was caused by …

  8. Focus-aware Response Generation in Inquiry Conversation

    2023

    Inquiry conversation is a common form of conversation that aims to complete the investigation (e.g., court hearing, medical consultation and police interrogation) during which a series of focus shifts occurs. While many models have been …

  9. MEDOE: A Multi-Expert Decoder and Output Ensemble Framework for Long-tailed Semantic Segmentation

    2023 · arXiv (Cornell University)

    Long-tailed distribution of semantic categories, which has been often ignored in conventional methods, causes unsatisfactory performance in semantic segmentation on tail categories. In this paper, we focus on the problem of long-tailed semantic segmentation. Although …

  10. Transferring Causal Mechanism over Meta-representations for Target-Unknown Cross-domain Recommendation

    2024 · ACM Transactions on Information Systems

    Tackling the pervasive issue of data sparsity in recommender systems, we present an insightful investigation into the burgeoning area of non-overlapping cross-domain recommendation, a technique that facilitates the transfer of interaction knowledge across domains without …

  11. ProfLLM: A framework for adapting offline large language models to few-shot expert knowledge

    2023

    Large language models perform well at common field, but are much less effective in niche academic fields .The cause of this problem is that large language models lack the ability to handle few-shot expert knowledge. …

  12. More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs

    2024 · arXiv (Cornell University)

    The performance on general tasks decreases after Large Language Models (LLMs) are fine-tuned on domain-specific tasks, the phenomenon is known as Catastrophic Forgetting (CF). However, this paper presents a further challenge for real application of …

  13. Application of Temporal Action Detection Technology in Abnormal Event Detection of Surveillance Video

    2025 · IEEE Access

    By detecting abnormal violation event in surveillance videos, the safety management capabilities in high-risk power operations can be improved. This research constructs an intelligent abnormal event detection technology using deep learning algorithms, aiming to improve …

  14. General information metrics for improving AI model training efficiency

    2025 · Artificial Intelligence Review

    Abstract To address the growing size of AI model training data and the lack of a universal data selection methodology–factors that significantly drive up training costs–this paper presents the General Information Metrics Evaluation (GIME) method. …

  15. Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

    2025 · arXiv (Cornell University)

    Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, …

  16. Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning

    2025 · arXiv (Cornell University)

    Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to their reliance on data-driven …

  17. Adaptive encrypted traffic classification via online hash center evolution

    2026 · Journal of King Saud University - Computer and Information Sciences

    With the widespread adoption of network encryption protocols, encrypted traffic enhances communication privacy while posing new challenges for threat detection. Although deep learning based encrypted traffic classification has achieved notable progress in recent years, most …

  18. Attentional Factorization Machines: Learning the Weight of Feature Interactions via Attention Networks

    2017

    Factorization Machines (FMs) are a supervised learning approach that enhances the linear regression model by incorporating the second-order feature interactions. Despite effectiveness, FM can be hindered by its modelling of all feature interactions with the …

  19. Dice Loss for Data-imbalanced NLP Tasks

    2020

    Many NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training. The …

  20. CorefQA: Coreference Resolution as Query-based Span Prediction

    2020

    In this paper, we present CorefQA, an accurate and extensible approach for the coreference resolution task. We formulate the problem as a span prediction task, like in question answering: A query is generated for each …

  21. A Unified MRC Framework for Named Entity Recognition

    2020

    The task of named entity recognition (NER) is normally divided into nested NER and flat NER depending on whether named entities are nested or not. Models are usually separately developed for the two tasks, since …