ملف الباحث

Pengjun Xie

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    Domain adaption for word segmentation and POS tagging is a challenging problem for Chinese lexical processing. Self-training is one promising solution for it, which struggles to construct a set of high-quality pseudo training instances for …

  2. Entity-to-Text based Data Augmentation for various Named Entity Recognition Tasks

    2022 · arXiv (Cornell University)

    Data augmentation techniques have been used to alleviate the problem of scarce labeled data in various NER tasks (flat, nested, and discontinuous NER tasks). Existing augmentation techniques either manipulate the words in the original text …

  3. COMBO: A Complete Benchmark for Open KG Canonicalization

    2023 · arXiv (Cornell University)

    Open knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text. The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need …

  4. Knowledge Mechanisms in Large Language Models: A Survey and Perspective

    2024

    Mengru Wang, Yunzhi Yao, Ziwen Xu, Shuofei Qiao, Shumin Deng, Peng Wang, Xiang Chen, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Findings of the Association for Computational Linguistics: EMNLP 2024. …

  5. SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement

    2025

    Runnan Fang, Xiaobin Wang, Yuan Liang, Shuofei Qiao, Jialong Wu, Zekun Xi, Ningyu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume …

  6. Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics

    2025 · arXiv (Cornell University)

    RAG (Retrieval-Augmented Generation) systems and web agents are increasingly evaluated on multi-hop deep search tasks, yet current practice suffers from two major limitations. First, most benchmarks leak the reasoning path in the question text, allowing …