Pith. sign in

REVIEW 2 cited by

Towards Effective Extraction and Evaluation of Factual Claims

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.10855 v2 pith:MYCDIMJ3 submitted 2025-02-15 cs.CL

classification cs.CL
keywords claimclaimsextractionframeworkevaluationfact-checkingmethodsclaimify
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A common strategy for fact-checking long-form content generated by Large Language Models (LLMs) is extracting simple claims that can be verified independently. Since inaccurate or incomplete claims compromise fact-checking results, ensuring claim quality is critical. However, the lack of a standardized evaluation framework impedes assessment and comparison of claim extraction methods. To address this gap, we propose a framework for evaluating claim extraction in the context of fact-checking along with automated, scalable, and replicable methods for applying this framework, including novel approaches for measuring coverage and decontextualization. We also introduce Claimify, an LLM-based claim extraction method, and demonstrate that it outperforms existing methods under our evaluation framework. A key feature of Claimify is its ability to handle ambiguity and extract claims only when there is high confidence in the correct interpretation of the source text.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new benchmark and 3D decomposition paradigm for factuality evaluation of interpretive claims about contact center conversations, with best LLM-judge F1 of 0.86.

  2. UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction

    cs.CL 2025-09 conditional novelty 3.0 of 10

    Fine-tuned FLAN-T5-Large achieved the top METEOR score for claim extraction, while prompting methods scored lower but produced claims that human reviewers found more useful in some cases.

Pith tools