article Open access

Structure-Aware Generative Information Extraction via Feature Space Alignment

  • Information
  • Multidisciplinary Digital Publishing Institute
Research footprint

At a glance

Citations
0
References
24
Comments
0
Paper overview

Abstract

Large language models (LLMs) face difficulties in leveraging the syntactic structures and entity relations embedded in text for long-document information extraction. To address this issue, this paper proposes a generative extraction method integrating heterogeneous topology awareness and spatial alignment. The method first extracts syntactic and coreference information to construct a heterogeneous document graph and employs a mixture-of-experts network to decouple and encode multi-type topological features. A component orthogonal projection mechanism and a graph-text contrastive learning strategy are then utilized to align the extracted graph features to the underlying semantic space of the language model with high fidelity. Furthermore, Topology-Aware Encoder compresses the global features into fixed-length structural prompts to guide text generation. Experiments on the ACE2005, WikiEvents, and DuEE datasets demonstrated that the proposed method achieved state-of-the-art performance on information extraction tasks. Consequently, these results suggest that the proposed framework is a promising approach for complex information extraction across base LLMs of different scales.

Record transparency

Publication details

DOI
10.3390/info17050409
OpenAlex
W7155551298
Document type
article
Language
EN
Source
Information
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.