Pith. sign in

REVIEW 1 cited by

Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08918 v1 pith:CKY2ZSIE submitted 2024-10-11 cs.CY

classification cs.CY
keywords wikimediacommunitydatagreatercontentdatasetslanguagemodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across pre-training, post-training, and model evaluations. We point to opportunities for greater use of Wikimedia content but also identify ways in which the language modeling community could better center the needs of Wikimedia editors. In particular, we call for incorporating additional sources of Wikimedia data, a greater focus on benchmarks for LLMs that encode Wikimedia principles, and greater multilingualism in Wikimedia-derived datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. If open source is to win, it must go public

    cs.CY 2025-07 conditional novelty 5.0 of 10

    Open source AI will not democratize access on its own, so the paper argues open models must be embedded in publicly funded and governed public AI infrastructure.

Pith tools