Researcher profile

Fandong Meng

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation Generation

    2021 · Proceedings of the AAAI Conference on Artificial Intelligence

    The success of emotional conversation systems depends on sufficient perception and appropriate expression of emotions. In a real-world conversation, we firstly instinctively perceive emotions from multi-source information, including the emotion flow of dialogue history, facial …

  2. Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation

    2021 · arXiv (Cornell University)

    Recently, knowledge distillation (KD) has shown great success in BERT compression. Instead of only learning from the teacher's soft label as in conventional KD, researchers find that the rich information contained in the hidden layers …

  3. Target-oriented Fine-tuning for Zero-Resource Named Entity Recognition

    2021

    Zero-resource named entity recognition (NER) severely suffers from data scarcity in a specific domain or language. Most studies on zero-resource NER transfer knowledge from various data by fine-tuning on different auxiliary tasks. However, how to …

  4. An Iterative Multi-Knowledge Transfer Network for Aspect-Based Sentiment Analysis

    2020 · arXiv (Cornell University)

    Aspect-based sentiment analysis (ABSA) mainly involves three subtasks: aspect term extraction, opinion term extraction, and aspect-level sentiment classification, which are typically handled in a separate or joint manner. However, previous approaches do not well exploit …

  5. Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information Extraction

    2023 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Zero-shot information extraction (IE) tasks have attracted great attention recently. However, how to jointly model multiple IE tasks in the zero-shot scenario is still an open question. In this article, we focus on zero-shot joint …

  6. LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation

    2024 · arXiv (Cornell University)

    Multilingual neural machine translation models generally distinguish translation directions by the language tag (LT) in front of the source or target sentences. However, current LT strategies cannot indicate the desired target language as expected on …

  7. Improving Machine Translation with Large Language Models: A Preliminary Study with Cooperative Decoding

    2024

    Contemporary translation engines based on the encoder-decoder framework have made significant strides in development.However, the emergence of Large Language Models (LLMs) has disrupted their position by presenting the potential for achieving superior translation quality.To uncover …

  8. Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific Loss

    2025

    Recent studies have explored Continual Instruction Tuning (CIT) in Multimodal Large Language Models (MLLMs), with a primary focus on Task-incremental CIT, where MLLMs are required to continuously acquire new tasks.However, the more practical and challenging …

  9. CM-Align: Consistency-based Multilingual Alignment for Large Language Models

    2025 · arXiv (Cornell University)

    Current large language models (LLMs) generally show a significant performance gap in alignment between English and other languages. To bridge this gap, existing research typically leverages the model's responses in English as a reference to …

  10. GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling

    2019

    Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs). However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information …

  11. Modeling Localness for Self-Attention Networks

    2018

    Self-attention networks have proven to be of profound value for its strength of capturing global dependencies. In this work, we propose to model localness for self-attention networks, which enhances the ability of capturing useful local …

  12. Bridging the Gap between Training and Inference for Neural Machine Translation

    2019

    Neural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words. At training time, it predicts with the ground truth words as context while at …

  13. Unsupervised Paraphrasing by Simulated Annealing

    2020

    We propose UPSA, a novel approach that accomplishes Unsupervised Paraphrasing by Simulated Annealing. We model paraphrase generation as an optimization problem and propose a sophisticated objective function, involving semantic similarity, expression diversity, and language fluency …

  14. Is ChatGPT a Good NLG Evaluator? A Preliminary Study

    2023

    Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation metrics. However, …