preprint Open access

Reference String Extraction Using Line-Based Conditional Random Fields

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
2
References
7
Comments
0
Paper overview

Öz

The extraction of individual reference strings from the reference section of scientific publications is an important step in the citation extraction pipeline. Current approaches divide this task into two steps by first detecting the reference section areas and then grouping the text lines in such areas into reference strings. We propose a classification model that considers every line in a publication as a potential part of a reference string. By applying line-based conditional random fields rather than constructing the graphical model based on the individual words, dependencies and patterns that are typical in reference sections provide strong features while the overall complexity of the model is reduced.

Record transparency

Publication details

DOI
10.48550/arxiv.1705.08154
OpenAlex
W2619751974
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.