Fei Wu
21 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Ad Recommendation for Sponsored Search Engine via Composite Long-Short Term Memory
2016
Search engine logs contain a large amount of users' click-through data that can be leveraged as implicit indicators of relevance. In this paper we address ad recommendation problem that finding and ranking the most relevant …
-
Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
2021
Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process …
-
Collaborative Semantic Aggregation and Calibration for Federated Domain Generalization
2021 · arXiv (Cornell University)
Domain generalization (DG) aims to learn from multiple known source domains a model that can generalize well to unknown target domains. The existing DG methods usually exploit the fusion of shared multi-source data to train …
-
A General Framework for Defending Against Backdoor Attacks via Influence Graph
2021 · arXiv (Cornell University)
In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attacking pattern, and therefore, poisoned training examples …
-
Combined Forecasting of Ship Heave Motion Based on Induced Ordered Weighted Averaging Operator
2022 · IEEJ Transactions on Electrical and Electronic Engineering
Abstract Heave motion of ships is a complex nonlinear dynamic process and cannot be accurately forecasted using a single prediction model. In this paper, an effective combined forecasting method is proposed to perform ship's heave …
-
GPT-NER: Named Entity Recognition via Large Language Models
2023 · arXiv (Cornell University)
Despite the fact that large-scale Language Models (LLM) have achieved SOTA performances on a variety of NLP tasks, its performance on NER is still significantly below supervised baselines. This is due to the gap between …
-
Pushing the Limits of ChatGPT on NLP Tasks
2023 · arXiv (Cornell University)
Despite the success of ChatGPT, its performances on most NLP tasks are still well below the supervised baselines. In this work, we looked into the causes, and discovered that its subpar performance was caused by …
-
Focus-aware Response Generation in Inquiry Conversation
2023
Inquiry conversation is a common form of conversation that aims to complete the investigation (e.g., court hearing, medical consultation and police interrogation) during which a series of focus shifts occurs. While many models have been …
-
MEDOE: A Multi-Expert Decoder and Output Ensemble Framework for Long-tailed Semantic Segmentation
2023 · arXiv (Cornell University)
Long-tailed distribution of semantic categories, which has been often ignored in conventional methods, causes unsatisfactory performance in semantic segmentation on tail categories. In this paper, we focus on the problem of long-tailed semantic segmentation. Although …
-
Transferring Causal Mechanism over Meta-representations for Target-Unknown Cross-domain Recommendation
2024 · ACM Transactions on Information Systems
Tackling the pervasive issue of data sparsity in recommender systems, we present an insightful investigation into the burgeoning area of non-overlapping cross-domain recommendation, a technique that facilitates the transfer of interaction knowledge across domains without …
-
ProfLLM: A framework for adapting offline large language models to few-shot expert knowledge
2023
Large language models perform well at common field, but are much less effective in niche academic fields .The cause of this problem is that large language models lack the ability to handle few-shot expert knowledge. …
-
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs
2024 · arXiv (Cornell University)
The performance on general tasks decreases after Large Language Models (LLMs) are fine-tuned on domain-specific tasks, the phenomenon is known as Catastrophic Forgetting (CF). However, this paper presents a further challenge for real application of …
-
Application of Temporal Action Detection Technology in Abnormal Event Detection of Surveillance Video
2025 · IEEE Access
By detecting abnormal violation event in surveillance videos, the safety management capabilities in high-risk power operations can be improved. This research constructs an intelligent abnormal event detection technology using deep learning algorithms, aiming to improve …
-
General information metrics for improving AI model training efficiency
2025 · Artificial Intelligence Review
Abstract To address the growing size of AI model training data and the lack of a universal data selection methodology–factors that significantly drive up training costs–this paper presents the General Information Metrics Evaluation (GIME) method. …
-
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
2025 · arXiv (Cornell University)
Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, …
-
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
2025 · arXiv (Cornell University)
Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to their reliance on data-driven …
-
Adaptive encrypted traffic classification via online hash center evolution
2026 · Journal of King Saud University - Computer and Information Sciences
With the widespread adoption of network encryption protocols, encrypted traffic enhances communication privacy while posing new challenges for threat detection. Although deep learning based encrypted traffic classification has achieved notable progress in recent years, most …
-
Attentional Factorization Machines: Learning the Weight of Feature Interactions via Attention Networks
2017
Factorization Machines (FMs) are a supervised learning approach that enhances the linear regression model by incorporating the second-order feature interactions. Despite effectiveness, FM can be hindered by its modelling of all feature interactions with the …
-
Dice Loss for Data-imbalanced NLP Tasks
2020
Many NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training. The …
-
CorefQA: Coreference Resolution as Query-based Span Prediction
2020
In this paper, we present CorefQA, an accurate and extensible approach for the coreference resolution task. We formulate the problem as a span prediction task, like in question answering: A query is generated for each …
-
A Unified MRC Framework for Named Entity Recognition
2020
The task of named entity recognition (NER) is normally divided into nested NER and flat NER depending on whether named entities are nested or not. Models are usually separately developed for the two tasks, since …