Pith. sign in

REVIEW 2 cited by

AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13117 v3 pith:G6YTHVMI submitted 2023-05-22 cs.CL

classification cs.CL
keywords evidenceclaimclaimsaveritecdatasetincludingreal-worldsubstantial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Existing datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we introduce AVeriTeC, a new dataset of 4,568 real-world claims covering fact-checks by 50 different organizations. Each claim is annotated with question-answer pairs supported by evidence available online, as well as textual justifications explaining how the evidence combines to produce a verdict. Through a multi-round annotation process, we avoid common pitfalls including context dependence, evidence insufficiency, and temporal leakage, and reach a substantial inter-annotator agreement of $\kappa=0.619$ on verdicts. We develop a baseline as well as an evaluation scheme for verifying claims through several question-answering steps against the open web.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 7 citations worldwide. Full citation record

  1. MEDIAREF: A Public Knowledge Store for Media Background Checks

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    MEDIAREF is a public, updatable web-document store that lets LLMs generate media background checks more reproducibly and with higher fact recall than zero-shot generation alone.

  2. Evidence-Ledger Adjudication for Claim-Evidence Traceability

    cs.AI 2026-07 conditional novelty 4.0 of 10

    An evidence-ledger workflow labels claim-evidence pairs as supported/contradicted/missing/mixed and routes unsupported claims back to authors, reporting 0.676 accuracy over TF-IDF's 0.383 on a 2,335-row benchmark.

Pith tools