Pith. sign in

REVIEW 1 cited by

Named Entity Recognition and Classification on Historical Documents: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.11406 v1 pith:WNQSW7H3 submitted 2021-09-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords historicaldocumentsnamedrecognitionclassificationentityopportunitiessurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

After decades of massive digitisation, an unprecedented amount of historical documents is available in digital format, along with their machine-readable texts. While this represents a major step forward with respect to preservation and accessibility, it also opens up new opportunities in terms of content mining and the next fundamental challenge is to develop appropriate technologies to efficiently search, retrieve and explore information from this 'big data of the past'. Among semantic indexing opportunities, the recognition and classification of named entities are in great demand among humanities scholars. Yet, named entity recognition (NER) systems are heavily challenged with diverse, historical and noisy inputs. In this survey, we present the array of challenges posed by historical documents to NER, inventory existing resources, describe the main approaches deployed so far, and identify key priorities for future developments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A fine-tuned GliNER model reaches 68.98% exact F1 and 75.64% fuzzy F1 on a new Italian historical NER benchmark, outperforming zero-shot LLaMa3.1-8B and zero-shot GliNER.

Pith tools