Researcher profile

Juanzi Li

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Entity Matching across Heterogeneous Sources

    2015

    Given an entity in a source domain, finding its matched entities from another (target) domain is an important task in many applications. Traditionally, the problem was usually addressed by first extracting major keywords corresponding to …

  2. Differentiating Concepts and Instances for Knowledge Graph Embedding

    2018 · ArXiv.org

    Concepts, which represent a group of different instances sharing common properties, are essential information in knowledge representation. Most conventional knowledge embedding methods encode both entities (concepts and instances) and relations as vectors in a low …

  3. MAVEN: A Massive General Domain Event Detection Dataset

    2020

    Xiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang, Rong Han, Zhiyuan Liu, Juanzi Li, Peng Li, Yankai Lin, Jie Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.

  4. KQA Pro: A Large-Scale Dataset with Interpretable Programs and Accurate SPARQLs for Complex Question Answering over Knowledge Base

    2020 · arXiv (Cornell University)

    Complex question answering over knowledge base (Complex KBQA) is challenging because it requires various compositional reasoning capabilities, such as multi-hop inference, attribute comparison, set operation, and etc. Existing benchmarks have some shortcomings that limit the …

  5. MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

    2022 · arXiv (Cornell University)

    The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due …

  6. Reasoning over Hierarchical Question Decomposition Tree for Explainable Question Answering

    2023 · arXiv (Cornell University)

    Explainable question answering (XQA) aims to answer a given question and provide an explanation why the answer is selected. Existing XQA methods focus on reasoning on a single knowledge source, e.g., structured knowledge bases, unstructured …

  7. Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach

    2023 · arXiv (Cornell University)

    Large Language Models (LLMs) have not only exhibited exceptional performance across various tasks, but also demonstrated sparks of intelligence. Recent studies have focused on assessing their capabilities on human exams and revealed their impressive competence …

  8. Explainable Few-shot Knowledge Tracing

    2024 · arXiv (Cornell University)

    Knowledge tracing (KT), aiming to mine students' mastery of knowledge by their exercise records and predict their performance on future test questions, is a critical task in educational assessment. While researchers achieved tremendous success with …

  9. LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA

    2024 · arXiv (Cornell University)

    Though current long-context large language models (LLMs) have demonstrated impressive capacities in answering user questions based on extensive text, the lack of citations in their responses makes user verification difficult, leading to concerns about their …

  10. Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models

    2025

    Large language models (LLMs) have been increasingly applied to various domains, which triggers increasing concerns about LLMs' safety on specialized domains, e.g. medicine. Despite prior explorations on general jailbreaking attacks, there are two challenges for …

  11. Text-enhanced representation learning for knowledge graph

    2016 · International Joint Conference on Artificial Intelligence

    Learning the representations of a knowledge graph has attracted significant research interest in the field of intelligent Web. By regarding each relation as one translation from head entity to tail entity, translation-based methods including TransE, …

  12. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2019 · arXiv (Cornell University)

    Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE …

  13. KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation

    2021 · Transactions of the Association for Computational Linguistics

    Abstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional …

  14. Parameter-efficient fine-tuning of large-scale pre-trained language models

    2023 · Nature Machine Intelligence

    Abstract With the prevalence of pre-trained language models (PLMs) and the pre-training–fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all …