Pith. sign in

REVIEW 2 cited by

Evaluating end-to-end entity linking on domain-specific knowledge bases: Learning about ancient technologies from museum collections

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14588 v1 pith:B3LGU3YA submitted 2023-05-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelapproachesdatadatasetdomaindomain-specificend-to-endentity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

To study social, economic, and historical questions, researchers in the social sciences and humanities have started to use increasingly large unstructured textual datasets. While recent advances in NLP provide many tools to efficiently process such data, most existing approaches rely on generic solutions whose performance and suitability for domain-specific tasks is not well understood. This work presents an attempt to bridge this domain gap by exploring the use of modern Entity Linking approaches for the enrichment of museum collection data. We collect a dataset comprising of more than 1700 texts annotated with 7,510 mention-entity pairs, evaluate some off-the-shelf solutions in detail using this dataset and finally fine-tune a recent end-to-end EL model on this data. We show that our fine-tuned model significantly outperforms other approaches currently available in this domain and present a proof-of-concept use case of this model. We release our dataset and our best model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enriching Social Science Research via Survey Item Linking

    cs.DL 2024-12 conditional novelty 6.0 of 10

    A new bilingual benchmark (SILD) and a two-stage pipeline make survey item linking feasible, reaching document-level recall above 50% at k=10, with mention detection as the main bottleneck.

  2. Unsupervised Named Entity Disambiguation for Low Resource Domains

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An unsupervised entity disambiguation method using Group Steiner Trees over knowledge graph subgraphs reports over 40% average improvement in Precision@1 over prior unsupervised baselines on four domain-specific datasets.

Pith tools