Structure-Aware Generative Information Extraction via Feature Space Alignment
At a glance
- Citations
- 0
- References
- 24
- Comments
- 0
Abstract
Large language models (LLMs) face difficulties in leveraging the syntactic structures and entity relations embedded in text for long-document information extraction. To address this issue, this paper proposes a generative extraction method integrating heterogeneous topology awareness and spatial alignment. The method first extracts syntactic and coreference information to construct a heterogeneous document graph and employs a mixture-of-experts network to decouple and encode multi-type topological features. A component orthogonal projection mechanism and a graph-text contrastive learning strategy are then utilized to align the extracted graph features to the underlying semantic space of the language model with high fidelity. Furthermore, Topology-Aware Encoder compresses the global features into fixed-length structural prompts to guide text generation. Experiments on the ACE2005, WikiEvents, and DuEE datasets demonstrated that the proposed method achieved state-of-the-art performance on information extraction tasks. Consequently, these results suggest that the proposed framework is a promising approach for complex information extraction across base LLMs of different scales.
Publication details
- DOI
- 10.3390/info17050409
- OpenAlex
- W7155551298
- Document type
- article
- Language
- EN
- Source
- Information
- Last metadata update
Comments
Log in to join the discussion.