Pith. sign in

REVIEW 9 cited by

Generative Large Language Models in Automated Fact-Checking: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02351 v3 pith:UAAKD236 submitted 2024-07-02 cs.CL

classification cs.CL
keywords fact-checkinggenerativellmsautomatedresearchsurveycurrentlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rapid spread of false and misleading information on online platforms poses a growing societal challenge, overwhelming the capacity of manual fact-checking and increasing the demand for scalable, reliable automation. Recent advances in generative large language models (LLMs) have broadened the scope of automated fact-checking beyond accuracy-driven prediction. LLMs are now integral components of fact-checking pipelines, supporting tasks such as generating new data, performing and assisting with fact verification, and shaping how fact-checking systems are evaluated. This survey provides a comprehensive overview of the role of generative LLMs in automated fact-checking, based on a systematic review of 199 research papers. We introduce a unifying taxonomy that captures how generative LLMs are integrated into fact-checking workflows and analyze their use across core fact-checking tasks, dataset construction and augmentation strategies, task formulations, and evaluation practices. Additionally, we investigate the impact of generative LLMs in multilingual and low-resource settings in fact-checking, highlighting trends, limitations, and gaps in current research. By consolidating fragmented research efforts and identifying methodological patterns, limitations, and open challenges, this survey maps the current state of generative LLMs in automated fact-checking. It aims to support researchers in developing more reliable, interpretable, and inclusive fact-checking systems, while outlining promising directions for future research in this rapidly evolving field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Attacks Against Automated Fact-Checking: A Survey

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A survey that organizes adversarial attacks on automated fact-checking into claim, evidence, and claim-evidence pair categories, and reports that current defenses address 13 of the 53 surveyed attacks.

  2. SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

    cs.IR 2026-07 conditional novelty 5.0 of 10

    A prompt-based multi-agent pipeline with an NLI citation auditor beats the CheckThat! 2026 Task 3 baseline on mean score (0.329 vs 0.272) by improving citation precision and recall, while trailing the baseline on enta...

  3. CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking

    cs.CL 2025-09 conditional novelty 5.0 of 10

    CANDY, a Chinese misinformation fact-checking benchmark, shows LLMs reach only ~76% accuracy on contamination-free claims and frequently fabricate supporting evidence, while serving better as human assistants than aut...

  4. RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new 6K-claim benchmark evaluates LLMs and multimodal LLMs on real-world fact-checking with an explicit 'unknown' option and shows web search and multimodal input improve performance.

  5. Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation

    cs.CL 2025-12 conditional novelty 4.0 of 10

    A deterministic, proposition-level fact-checker that compares clinical summaries against electronic health records via (entity, attribute, value, time) claims and hard-coded logical checks reports 0.8904 precision and...

  6. DS@GT at CheckThat! 2025: Ensemble Methods for Detection of Scientific Discourse on Social Media

    cs.CL 2025-07 conditional novelty 4.0 of 10

    The DS@GT system achieved 0.8611 macro-F1 on the CheckThat! 2025 Task 4a development set by combining a fine-tuned DeBERTa model with GPT-4o few-shot prompting, outperforming the DeBERTaV3 baseline of 0.8375.

  7. Leveraging Large Language Models for Information Verification -- an Engineering Approach

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A GPT-4o based pipeline that searches the web, picks keyframes, transcribes audio, and cross-validates everything to produce news verification reports.

  8. A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages

    cs.CL 2025-04 conditional novelty 4.0 of 10

    An LLM-based multilingual pipeline retrieves previously fact-checked claims, filters irrelevant ones, summarizes articles, and predicts veracity, evaluated across 23 languages.

  9. Challenges in Guardrailing Large Language Models for Science

    cs.AI 2024-11 conditional novelty 3.0 of 10

    A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.

Pith tools