Researcher profile

Xiaozhi Wang

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. MAVEN: A Massive General Domain Event Detection Dataset

    2020

    Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.

  2. SHUOWEN-JIEZI: Linguistically Informed Tokenizers For Chinese Language Model Pretraining

    2021 · arXiv (Cornell University)

    Conventional tokenization methods for Chinese pretrained language models (PLMs) treat each character as an indivisible token (Devlin et al., 2019), which ignores the characteristics of the Chinese writing system. In this work, we comprehensively study …

  3. MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

    2022 · arXiv (Cornell University)

    The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due …

  4. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2019 · arXiv (Cornell University)

    Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE …

  5. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2021 · Transactions of the Association for Computational Linguistics

    Abstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional …

  6. Parameter-efficient fine-tuning of large-scale pre-trained language models

    2023 · Nature Machine Intelligence

    Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …