Xiaozhi Wang
6 papers in the PaperMetrix corpus
Papers by this author
-
MAVEN: A Massive General Domain Event Detection Dataset
2020
Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
-
SHUOWEN-JIEZI: Linguistically Informed Tokenizers For Chinese Language Model Pretraining
2021 · arXiv (Cornell University)
Conventional tokenization methods for Chinese pretrained language models (PLMs) treat each character as an indivisible token (Devlin et al., 2019), which ignores the characteristics of the Chinese writing system. In this work, we comprehensively study …
-
MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction
2022 · arXiv (Cornell University)
The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due …
-
KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
2019 · arXiv (Cornell University)
Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE …
-
KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
2021 · Transactions of the Association for Computational Linguistics
Abstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional …
-
Parameter-efficient fine-tuning of large-scale pre-trained language models
2023 · Nature Machine Intelligence
Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …