Researcher profile

Zhiyuan Liu

45 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Learning Cross-lingual Word Embeddings via Matrix Co-factorization

    2015

    Tianze Shi, Zhiyuan Liu, Yang Liu, Maosong Sun. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 2015.

  2. Block Belief Propagation for Parameter Learning in Markov Random Fields

    2018 · arXiv (Cornell University)

    Traditional learning methods for training Markov random fields require doing inference over all variables to compute the likelihood gradient. The iteration complexity for those methods therefore scales with the size of the graphical models. In …

  3. Differentiating Concepts and Instances for Knowledge Graph Embedding

    2018 · ArXiv.org

    Concepts, which represent a group of different instances sharing common properties, are essential information in knowledge representation. Most conventional knowledge embedding methods encode both entities (concepts and instances) and relations as vectors in a low …

  4. Country Image in COVID-19 Pandemic: A Case Study of China

    2020 · IEEE Transactions on Big Data

    Country image has a profound influence on international relations and economic development. In the worldwide outbreak of COVID-19, countries and their people display different reactions, resulting in diverse perceived images among foreign public. Therefore, in …

  5. OpenAttack: An Open-source Textual Adversarial Attack Toolkit

    2021

    Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, Maosong Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint …

  6. Denoising Relation Extraction from Document-level Distant Supervision

    2020

    Distant supervision (DS) has been widely used to generate auto-labeled data for sentencelevel relation extraction (RE), which improves RE performance. However, the existing success of DS cannot be directly transferred to the more challenging document-level …

  7. Learning from Context or Names? An Empirical Study on Neural Relation Extraction

    2020

    Neural models have achieved remarkable success on relation extraction (RE) benchmarks. However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance …

  8. MAVEN: A Massive General Domain Event Detection Dataset

    2020

    Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.

  9. SHUOWEN-JIEZI: Linguistically Informed Tokenizers For Chinese Language Model Pretraining

    2021 · arXiv (Cornell University)

    Conventional tokenization methods for Chinese pretrained language models (PLMs) treat each character as an indivisible token (Devlin et al., 2019), which ignores the characteristics of the Chinese writing system. In this work, we comprehensively study …

  10. Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    Adversarial attacks and backdoor attacks are two common security threats that hang over deep learning. Both of them harness taskirrelevant features of data in their implementation. Text style is a feature that is naturally irrelevant …

  11. MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

    2022 · arXiv (Cornell University)

    The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due …

  12. Rethinking Dense Retrieval's Few-Shot Ability

    2023 · arXiv (Cornell University)

    Few-shot dense retrieval (DR) aims to effectively generalize to novel search scenarios by learning a few samples. Despite its importance, there is little study on specialized datasets and standardized evaluation protocols. As a result, current …

  13. Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

    2023 · arXiv (Cornell University)

    Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading …

  14. Plug-and-Play Knowledge Injection for Pre-trained Language Models

    2023 · arXiv (Cornell University)

    Injecting external knowledge can improve the performance of pre-trained language models (PLMs) on various downstream NLP tasks. However, massive retraining is required to deploy new knowledge injection methods or knowledge bases for downstream tasks. In …

  15. Exploring Mode Connectivity for Pre-trained Language Models

    2022

    Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found. Although plenty of …

  16. Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub

    2023 · arXiv (Cornell University)

    Large Language Models (LLMs) excel in traditional natural language processing tasks but struggle with problems that require complex domain-specific calculations or simulations. While equipping LLMs with external tools to build LLM-based agents can enhance their …

  17. UniMem: Towards a Unified View of Long-Context Large Language Models

    2024 · arXiv (Cornell University)

    Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-context processing ability of LLMs, they are developed in an isolated …

  18. Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication

    2024 · arXiv (Cornell University)

    Natural language (NL) has long been the predominant format for human cognition and communication, and by extension, has been similarly pivotal in the development and application of Large Language Models (LLMs). Yet, besides NL, LLMs …

  19. PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes

    2024 · arXiv (Cornell University)

    Multimodal Large Language Models (MLLMs) have seen growing adoption across various scientific disciplines. These advancements encourage the investigation of molecule-text modeling within synthetic chemistry, a field dedicated to designing and conducting chemical reactions to synthesize …

  20. Modeling Relation Paths for Representation Learning of Knowledge Bases

    2015

    Representation learning of knowledge bases aims to embed both entities and relations into a low-dimensional space. Most existing methods only consider direct relations in representation learning. We argue that multiple-step relation paths also contain rich …

  21. Topical Word Embeddings

    2015 · Proceedings of the AAAI Conference on Artificial Intelligence

    Most word embedding models typically represent each word using a single vector, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance discriminativeness, we employ latent topic models to assign topics …

  22. A C-LSTM Neural Network for Text Classification

    2015 · arXiv (Cornell University)

    Neural network models have been demonstrated to be capable of achieving remarkable performance in sentence and document modeling. Convolutional neural network (CNN) and recurrent neural network (RNN) are two mainstream architectures for such modeling tasks, …

  23. Joint learning of character and word embeddings

    2015 · International Conference on Artificial Intelligence

    Most word embedding methods take a word as a basic unit and learn embeddings according to words' external contexts, ignoring the internal structures of words. However, in some languages such as Chinese, a word is …

  24. Relation Classification via Multi-Level Attention CNNs

    2016

    Relation classification is a crucial ingredient in numerous information extraction systems seeking to mine structured facts from text. We propose a novel convolutional neural network architecture for this task, relying on two levels of attention …

  25. Neural Relation Extraction with Selective Attention over Instances

    2016

    Distant supervised relation extraction has been widely used to find novel relational facts from text. However, distant supervision inevitably accompanies with the wrong labelling problem, and these noisy data will substantially hurt the performance of …

  26. End-to-End Neural Ad-hoc Ranking with Kernel Pooling

    2017

    This paper proposes K-NRM, a kernel based neural model for document ranking. Given a query and a set of documents, K-NRM uses a translation matrix that models word-level similarities via word embeddings, a new kernel-pooling …

  27. Neural Relation Extraction with Multi-lingual Attention

    2017

    Relation extraction has been widely used for finding unknown relational facts from the plain text. Most existing methods focus on exploiting mono-lingual data for relation extraction, ignoring massive information from the texts in various languages. …

  28. Convolutional Neural Networks for Soft-Matching N-Grams in Ad-hoc Search

    2018

    This paper presents \textttConv-KNRM, a Convolutional Kernel-based Neural Ranking Model that models n-gram soft matches for ad-hoc search. Instead of exact matching query and document n-grams, \textttConv-KNRM uses Convolutional Neural Networks to represent n-grams of …

  29. Neural Knowledge Acquisition via Mutual Attention Between Knowledge Graph and Text

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    We propose a general joint representation learning framework for knowledge acquisition (KA) on two tasks, knowledge graph completion (KGC) and relation extraction (RE) from text. In this framework, we learn representations of knowledge graphs (KGs) …

  30. Denoising Distantly Supervised Open-Domain Question Answering

    2018

    Distantly supervised open-domain question answering (DS-QA) aims to find answers in collections of unlabeled text. Existing DS-QA models usually retrieve related paragraphs from a large-scale corpus and apply reading comprehension technique to extract answers from …

  31. GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification

    2019

    Fact verification (FV) is a challenging task which requires to retrieve relevant evidence from plain text and use the evidence to verify given claims. Many claims require to simultaneously integrate and reason over several pieces …

  32. DocRED: A Large-Scale Document-Level Relation Extraction Dataset

    2019

    Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.

  33. ERNIE: Enhanced Language Representation with Informative Entities

    2019

    Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks. However, the existing …

  34. Incorporating Relation Paths in Neural Relation Extraction

    2017

    Distantly supervised relation extraction has been widely used to find novel relational facts from plain text. To predict the relation between a pair of two target entities, existing methods solely rely on those direct sentences …

  35. FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation

    2018

    We present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. The relation of each sentence is first recognized by distant supervision methods, …

  36. OpenNRE: An Open and Extensible Toolkit for Neural Relation Extraction

    2019

    Xu Han, Tianyu Gao, Yuan Yao, Deming Ye, Zhiyuan Liu, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): …

  37. Low-Resource Name Tagging Learned with Weakly Labeled Data

    2019

    Yixin Cao, Zikun Hu, Tat-seng Chua, Zhiyuan Liu, Heng Ji. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  38. FewRel 2.0: Towards More Challenging Few-Shot Relation Classification

    2019

    Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language …

  39. Multi-Interest Network with Dynamic Routing for Recommendation at Tmall

    2019

    Industrial recommender systems have embraced deep learning algorithms for building intelligent systems to make accurate recommendations. At its core, deep learning offers powerful ability for learning representations from data, especially for user and item representations. …

  40. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2019 · arXiv (Cornell University)

    Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE …

  41. Expertise Style Transfer: A New Task Towards Better Communication between Experts and Laymen

    2020

    The curse of knowledge can impede communication between experts and laymen. We propose a new task of expertise style transfer and contribute a manually annotated dataset with the goal of alleviating such cognitive biases. Solving …

  42. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2021 · Transactions of the Association for Computational Linguistics

    Abstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional …

  43. PTR: Prompt Tuning with Rules for Text Classification

    2022 · AI Open

    Recently, prompt tuning has been widely applied to stimulate the rich knowledge in pre-trained language models (PLMs) to serve NLP tasks. Although prompt tuning has achieved promising results on some few-class classification tasks, such as …

  44. Parameter-efficient fine-tuning of large-scale pre-trained language models

    2023 · Nature Machine Intelligence

    Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …