Pith. sign in

REVIEW

Structured references from PDF articles: assessing the tools for bibliographic reference extraction and parsing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14677 v2 pith:BXQKTSB7 submitted 2022-05-29 cs.DL

classification cs.DL
keywords toolsbibliographicreferencesanystyleareasarticlescermineextract
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many solutions have been provided to extract bibliographic references from PDF papers. Machine learning, rule-based and regular expressions approaches were among the most used methods adopted in tools for addressing this task. This work aims to identify and evaluate all and only the tools which, given a full-text paper in PDF format, can recognise, extract and parse bibliographic references. We identified seven tools: Anystyle, Cermine, ExCite, Grobid, Pdfssa4met, Scholarcy and Science Parse. We compared and evaluated them against a corpus of 56 PDF articles published in 27 subject areas. Indeed, Anystyle obtained the best overall score, followed by Cermine. However, in some subject areas, other tools had better results for specific tasks.

Discussion (0). Continue with ORCID to comment.

Pith tools