Pith. sign in

REVIEW 3 cited by

QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17169 v3 pith:L2OA3VUF submitted 2024-03-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords claimsnumericalexistingquantempreal-worldclaimdiverseevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated fact checking has gained immense interest to tackle the growing misinformation in the digital era. Existing systems primarily focus on synthetic claims on Wikipedia, and noteworthy progress has also been made on real-world claims. In this work, we release QuanTemp, a diverse, multi-domain dataset focused exclusively on numerical claims, encompassing temporal, statistical and diverse aspects with fine-grained metadata and an evidence collection without leakage. This addresses the challenge of verifying real-world numerical claims, which are complex and often lack precise information, not addressed by existing works that mainly focus on synthetic claims. We evaluate and quantify the limitations of existing solutions for the task of verifying numerical claims. We also evaluate claim decomposition based methods, numerical understanding based models and our best baselines achieves a macro-F1 of 58.32. This demonstrates that QuanTemp serves as a challenging evaluation set for numerical claim verification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Typed quantity verification exposes a canonical-equivalence blind spot in neural fact-checkers; Symbolic Augmentation fixes it (36.5%→98.2%) and transfers to SciFact-Open (+0.037 binary macro-F1).

  2. Sample Efficient Demonstration Selection for In-Context Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    CASE is a top-m linear bandit algorithm with challenger-arm sampling that selects exemplar subsets for in-context learning using up to 7x fewer LLM calls than prior methods.

  3. DS@GT at CheckThat! 2025: Evaluating Context and Tokenization Strategies for Numerical Fact Verification

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Longer context windows and right-to-left number tokenization do not improve numerical fact verification; evidence quality is the main bottleneck.

Pith tools