Researcher profile

Jie Zhou

38 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Data-driven Learning in Second Language Writing Class: A Survey of Empirical Studies

    2017 · International Journal of Emerging Technologies in Learning (iJET)

    Corpus technology is commonly used by researchers and professionals for language description; however it can also be employed by second or foreign language learners in what has come to be known as data-driven learning (DDL). …

  2. Deep Metric Learning via Adaptive Learnable Assessment

    2020

    In this paper, we propose a deep metric learning via adaptive learnable assessment (DML-ALA) method for image retrieval and clustering, which aims to learn a sample assessment strategy to maximize the generalization of the trained …

  3. Disentangle-based Continual Graph Representation Learning

    2020 · arXiv (Cornell University)

    Graph embedding (GE) methods embed nodes (and/or edges) in graph into a low-dimensional semantic space, and have shown its effectiveness in modeling multi-relational data. However, existing GE models are not practical in real-world applications since …

  4. Learning from Context or Names? An Empirical Study on Neural Relation Extraction

    2020

    Neural models have achieved remarkable success on relation extraction (RE) benchmarks. However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance …

  5. MAVEN: A Massive General Domain Event Detection Dataset

    2020

    Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.

  6. Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation Generation

    2021 · Proceedings of the AAAI Conference on Artificial Intelligence

    The success of emotional conversation systems depends on sufficient perception and appropriate expression of emotions. In a real-world conversation, we firstly instinctively perceive emotions from multi-source information, including the emotion flow of dialogue history, facial …

  7. Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation

    2021 · arXiv (Cornell University)

    Recently, knowledge distillation (KD) has shown great success in BERT compression. Instead of only learning from the teacher's soft label as in conventional KD, researchers find that the rich information contained in the hidden layers …

  8. Target-oriented Fine-tuning for Zero-Resource Named Entity Recognition

    2021

    Zero-resource named entity recognition (NER) severely suffers from data scarcity in a specific domain or language. Most studies on zero-resource NER transfer knowledge from various data by fine-tuning on different auxiliary tasks. However, how to …

  9. An Iterative Multi-Knowledge Transfer Network for Aspect-Based Sentiment Analysis

    2020 · arXiv (Cornell University)

    Aspect-based sentiment analysis (ABSA) mainly involves three subtasks: aspect term extraction, opinion term extraction, and aspect-level sentiment classification, which are typically handled in a separate or joint manner. However, previous approaches do not well exploit …

  10. OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions

    2022 · arXiv (Cornell University)

    The pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning. However, with the availability of massive labeled data, a natural question emerges: how …

  11. MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

    2022 · arXiv (Cornell University)

    The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due …

  12. Multimodal Aspect-Level Sentiment Analysis based on Deep Neural Networks

    2022

    Aspect-level sentiment analysis is a fine-grained task of sentiment analysis that aims to identify the sentiment polarity of specific aspect words in a sentence. However, most existing approaches rely mainly on text content and ignore …

  13. A Novel Radio Frequency Fingerprint Identification Method Using Incremental Learning

    2022 · 2022 IEEE 96th Vehicular Technology Conference (VTC2022-Fall)

    Radio frequency fingerprint (RFF) is regarded as a key technology in physical layer security in various wireless communications systems. Deep learning (DL) has achieved great success in the field of signal identification, particularly in improving …

  14. Feature Decomposition for Reducing Negative Transfer: A Novel Multi-task Learning Method for Recommender System

    2023 · arXiv (Cornell University)

    In recent years, thanks to the rapid development of deep learning (DL), DL-based multi-task learning (MTL) has made significant progress, and it has been successfully applied to recommendation systems (RS). However, in a recommender system, …

  15. Active Temporal Knowledge Graph Alignment

    2023 · International Journal on Semantic Web and Information Systems

    Entity alignment aims to identify equivalent entity pairs from different knowledge graphs (KGs). Recently, aligning temporal knowledge graphs (TKGs) that contain time information has aroused increasingly more interest, as the time dimension is widely used …

  16. Plug-and-Play Knowledge Injection for Pre-trained Language Models

    2023 · arXiv (Cornell University)

    Injecting external knowledge can improve the performance of pre-trained language models (PLMs) on various downstream NLP tasks. However, massive retraining is required to deploy new knowledge injection methods or knowledge bases for downstream tasks. In …

  17. Attacking Pre-trained Recommendation

    2023

    Recently, a series of pioneer studies have shown the potency of pre-trained models in sequential recommendation, illuminating the path of building an omniscient unified pre-trained recommendation model for different downstream recommendation tasks. Despite these advancements, …

  18. Mixture of Attention Heads: Selecting Attention Heads Per Token

    2022

    Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly focused on the feedforward layer in Transformer architecture. This …

  19. Exploring Mode Connectivity for Pre-trained Language Models

    2022

    Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found. Although plenty of …

  20. Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information Extraction

    2023 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Zero-shot information extraction (IE) tasks have attracted great attention recently. However, how to jointly model multiple IE tasks in the zero-shot scenario is still an open question. In this article, we focus on zero-shot joint …

  21. CausalABSC: Causal Inference for Aspect Debiasing in Aspect-Based Sentiment Classification

    2023 · IEEE/ACM Transactions on Audio Speech and Language Processing

    As the primary subtask of sentiment analysis, aspect-based sentiment classification (ABSC) aims to predict the sentiment polarity for a given aspect. While recent deep neural models for ABSC have shown good performance, their robustness is …

  22. OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions

    2023

    The pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning. However, with the availability of massive labeled data, a natural question emerges: how …

  23. LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation

    2024 · arXiv (Cornell University)

    Multilingual neural machine translation models generally distinguish translation directions by the language tag (LT) in front of the source or target sentences. However, current LT strategies cannot indicate the desired target language as expected on …

  24. Improving Machine Translation with Large Language Models: A Preliminary Study with Cooperative Decoding

    2024

    Contemporary translation engines based on the encoder-decoder framework have made significant strides in development.However, the emergence of Large Language Models (LLMs) has disrupted their position by presenting the potential for achieving superior translation quality.To uncover …

  25. Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific Loss

    2025

    Recent studies have explored Continual Instruction Tuning (CIT) in Multimodal Large Language Models (MLLMs), with a primary focus on Task-incremental CIT, where MLLMs are required to continuously acquire new tasks.However, the more practical and challenging …

  26. Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning

    2025

    Recent advancements in large language models (LLMs) have sparked significant interest in the automatic generation of Register Transfer Level (RTL) designs, particularly using Verilog. Current research on this topic primarily focuses on pre-training and instruction …

  27. Integrating Gaussian Process and C-Mixup for Regression

    2025

    This paper presents a novel approach that integrates Gaussian Process Regression (GPR) with C-Mixup, aiming to explore the synergistic potential of these two techniques in regression tasks. The GPR offers a flexible and powerful probabilistic …

  28. CM-Align: Consistency-based Multilingual Alignment for Large Language Models

    2025 · arXiv (Cornell University)

    Current large language models (LLMs) generally show a significant performance gap in alignment between English and other languages. To bridge this gap, existing research typically leverages the model's responses in English as a reference to …

  29. Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback

    2025 · arXiv (Cornell University)

    This paper introduces an interactive continual learning paradigm where AI models dynamically learn new skills from real-time human feedback while retaining prior knowledge. This paradigm distinctively addresses two major limitations of traditional continual learning: (1) …

  30. A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual Sets

    2025

    Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method …

  31. End-to-end learning of semantic role labeling using recurrent neural networks

    2015

    Jie Zhou, Wei Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  32. GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification

    2019

    Fact verification (FV) is a challenging task which requires to retrieve relevant evidence from plain text and use the evidence to verify given claims. Many claims require to simultaneously integrate and reason over several pieces …

  33. DocRED: A Large-Scale Document-Level Relation Extraction Dataset

    2019

    Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.

  34. GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling

    2019

    Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs). However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information …

  35. Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation

    2016 · Transactions of the Association for Computational Linguistics

    Neural machine translation (NMT) aims at solving machine translation (MT) problems using neural networks and has exhibited promising results in recent years. However, most of the existing NMT models are shallow and there is still …

  36. FewRel 2.0: Towards More Challenging Few-Shot Relation Classification

    2019

    Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …

  37. Unsupervised Paraphrasing by Simulated Annealing

    2020

    We propose UPSA, a novel approach that accomplishes Unsupervised Paraphrasing by Simulated Annealing. We model paraphrase generation as an optimization problem and propose a sophisticated objective function, involving semantic similarity, expression diversity, and language fluency …

  38. Is ChatGPT a Good NLG Evaluator? A Preliminary Study

    2023

    Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation metrics. However, …