Information Extraction and Knowledge Graph Construction of Chinese Scientific and Technical Literature Based on Fine-Tuning Large Language Models
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Öz
The effectiveness of literature retrieval depends on researchers’ experience and understanding of knowledge. To improve the retrieval efficiency and effect, this paper establishes an information extraction model based on a fine-tuned large language model for Chinese scientific and technological literature. It extracts specific information in the abstracts according to the readers' needs. The article manually labeled 1051 Chinese scientific and technical literature abstracts with the theme of ‘Knowledge Graph’, and constructed the experimental dataset of the model. The experiments compare the large language models such as DeepSeek, Llama3, Mistral, Gemma, etc. Meanwhile, the UIE model and the Bert model are compared with the large language models. The experiments show that the large language models have more obvious advantages in literature information extraction, and the F1 value of Mistral fine-tuning model validation ROUGE-L reaches 0.9213, which is about 0.21 higher than the un-fine-tuning F1 value, and 0.13 higher than that of UIE. Finally, this paper utilizes Neo4j to present the results of literature information extraction to the readers in the form of a knowledge graph.
Publication details
- DOI
- 10.1109/ecis65594.2025.11086863
- OpenAlex
- W4413158854
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.