ملف الباحث

Maosong Sun

41 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Learning Cross-lingual Word Embeddings via Matrix Co-factorization

    2015

    Tianze Shi, Zhiyuan Liu, Yang Liu, Maosong Sun. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.

  2. Generalized Agreement for Bidirectional Word Alignment

    2015

    While agreement-based joint training has proven to deliver state-of-the-art alignment accuracy, the produced word alignments are usually restricted to one-toone mappings because of the hard constraint on agreement. We propose a general framework to allow …

  3. Country Image in COVID-19 Pandemic: A Case Study of China

    2020 · IEEE Transactions on Big Data

    Country image has a profound influence on international relations and economic development. In the worldwide outbreak of COVID-19, countries and their people display different reactions, resulting in diverse perceived images among foreign public. Therefore, in …

  4. OpenAttack: An Open-source Textual Adversarial Attack Toolkit

    2021

    Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, Maosong Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint …

  5. Denoising Relation Extraction from Document-level Distant Supervision

    2020

    Distant supervision (DS) has been widely used to generate auto-labeled data for sentencelevel relation extraction (RE), which improves RE performance. However, the existing success of DS cannot be directly transferred to the more challenging document-level …

  6. Learning from Context or Names? An Empirical Study on Neural Relation Extraction

    2020

    Neural models have achieved remarkable success on relation extraction (RE) benchmarks. However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance …

  7. Neural Machine Translation With Explicit Phrase Alignment

    2021 · IEEE/ACM Transactions on Audio Speech and Language Processing

    While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads …

  8. SHUOWEN-JIEZI: Linguistically Informed Tokenizers For Chinese Language Model Pretraining

    2021 · arXiv (Cornell University)

    Conventional tokenization methods for Chinese pretrained language models (PLMs) treat each character as an indivisible token (Devlin et al., 2019), which ignores the characteristics of the Chinese writing system. In this work, we comprehensively study …

  9. Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    Adversarial attacks and backdoor attacks are two common security threats that hang over deep learning. Both of them harness taskirrelevant features of data in their implementation. Text style is a feature that is naturally irrelevant …

  10. A Template-based Method for Constrained Neural Machine Translation

    2022 · arXiv (Cornell University)

    Machine translation systems are expected to cope with various types of constraints in many practical scenarios. While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints …

  11. Semi-Supervised Learning for Neural Machine Translation

    2016 · arXiv (Cornell University)

    While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …

  12. Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

    2023 · arXiv (Cornell University)

    Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading …

  13. Plug-and-Play Knowledge Injection for Pre-trained Language Models

    2023 · arXiv (Cornell University)

    Injecting external knowledge can improve the performance of pre-trained language models (PLMs) on various downstream NLP tasks. However, massive retraining is required to deploy new knowledge injection methods or knowledge bases for downstream tasks. In …

  14. Exploring Mode Connectivity for Pre-trained Language Models

    2022

    Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found. Although plenty of …

  15. Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub

    2023 · arXiv (Cornell University)

    Large Language Models (LLMs) excel in traditional natural language processing tasks but struggle with problems that require complex domain-specific calculations or simulations. While equipping LLMs with external tools to build LLM-based agents can enhance their …

  16. UniMem: Towards a Unified View of Long-Context Large Language Models

    2024 · arXiv (Cornell University)

    Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-context processing ability of LLMs, they are developed in an isolated …

  17. Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication

    2024 · arXiv (Cornell University)

    Natural language (NL) has long been the predominant format for human cognition and communication, and by extension, has been similarly pivotal in the development and application of Large Language Models (LLMs). Yet, besides NL, LLMs …

  18. Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation

    2026

    Multimodal Retrieval-Augmented Generation (MRAG) has shown promise in mitigating hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge. However, existing methods typically adhere to rigid retrieval paradigms by mimicking fixed retrieval trajectories and …

  19. Modeling Relation Paths for Representation Learning of Knowledge Bases

    2015

    Representation learning of knowledge bases aims to embed both entities and relations into a low-dimensional space. Most existing methods only consider direct relations in representation learning. We argue that multiple-step relation paths also contain rich …

  20. Topical Word Embeddings

    2015 · Proceedings of the AAAI Conference on Artificial Intelligence

    Most word embedding models typically represent each word using a single vector, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance discriminativeness, we employ latent topic models to assign topics …

  21. Joint learning of character and word embeddings

    2015 · International Conference on Artificial Intelligence

    Most word embedding methods take a word as a basic unit and learn embeddings according to words' external contexts, ignoring the internal structures of words. However, in some languages such as Chinese, a word is …

  22. Neural Relation Extraction with Selective Attention over Instances

    2016

    Distant supervised relation extraction has been widely used to find novel relational facts from text. However, distant supervision inevitably accompanies with the wrong labelling problem, and these noisy data will substantially hurt the performance of …

  23. THUMT: An Open Source Toolkit for Neural Machine Translation

    2017 · arXiv (Cornell University)

    This paper introduces THUMT, an open-source toolkit for neural machine translation (NMT) developed by the Natural Language Processing Group at Tsinghua University. THUMT implements the standard attention-based encoder-decoder framework on top of Theano and supports …

  24. Neural Relation Extraction with Multi-lingual Attention

    2017

    Relation extraction has been widely used for finding unknown relational facts from the plain text. Most existing methods focus on exploiting mono-lingual data for relation extraction, ignoring massive information from the texts in various languages. …

  25. Adversarial Training for Unsupervised Bilingual Lexicon Induction

    2017

    Word embeddings are well known to capture linguistic regularities of the language on which they are trained. Researchers also observe that these regularities can transfer across languages. However, previous endeavors to connect separate monolingual word …

  26. Visualizing and Understanding Neural Machine Translation

    2017

    While neural machine translation (NMT) has made remarkable progress in recent years, it is hard to interpret its internal workings due to the continuous representations and non-linearity of neural networks. In this work, we propose …

  27. Neural Knowledge Acquisition via Mutual Attention Between Knowledge Graph and Text

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    We propose a general joint representation learning framework for knowledge acquisition (KA) on two tasks, knowledge graph completion (KGC) and relation extraction (RE) from text. In this framework, we learn representations of knowledge graphs (KGs) …

  28. Denoising Distantly Supervised Open-Domain Question Answering

    2018

    Distantly supervised open-domain question answering (DS-QA) aims to find answers in collections of unlabeled text. Existing DS-QA models usually retrieve related paragraphs from a large-scale corpus and apply reading comprehension technique to extract answers from …

  29. GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification

    2019

    Fact verification (FV) is a challenging task which requires to retrieve relevant evidence from plain text and use the evidence to verify given claims. Many claims require to simultaneously integrate and reason over several pieces …

  30. Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach

    2019

    While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors. In this work, we propose a contrastive learning approach to reducing word omission errors in NMT. The …

  31. DocRED: A Large-Scale Document-Level Relation Extraction Dataset

    2019

    Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.

  32. ERNIE: Enhanced Language Representation with Informative Entities

    2019

    Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks. However, the existing …

  33. Improving the Transformer Translation Model with Document-Level Context

    2018

    Although the Transformer translation model In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder. As large-scale document-level …

  34. Incorporating Relation Paths in Neural Relation Extraction

    2017

    Distantly supervised relation extraction has been widely used to find novel relational facts from plain text. To predict the relation between a pair of two target entities, existing methods solely rely on those direct sentences …

  35. Minimum Risk Training for Neural Machine Translation

    2016

    We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily …

  36. FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

    2018

    We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation of each sentence is first recognized by distant supervision methods, …

  37. OpenNRE: An Open and Extensible Toolkit for Neural Relation Extraction

    2019

    Xu Han, Tianyu Gao, Yuan Yao, Deming Ye, Zhiyuan Liu, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): …

  38. FewRel 2.0: Towards More Challenging Few-Shot Relation Classification

    2019

    Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …

  39. PTR: Prompt Tuning with Rules for Text Classification

    2022 · AI Open

    Recently, prompt tuning has been widely applied to stimulate the rich knowledge in pre-trained language models (PLMs) to serve NLP tasks. Although prompt tuning has achieved promising results on some few-class classification tasks, such as …

  40. Parameter-efficient fine-tuning of large-scale pre-trained language models

    2023 · Nature Machine Intelligence

    Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …