Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Recon, Answer, Verify: Agents in Search of Truth

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A cue-free fact-checking benchmark drops LLM macro-F1 by about 22 percent, showing how much models rely on leaked verdicts.

desk verdict Useful pipeline, overstated benchmark: the RAV agentic design is solid, but PFO's cleanup doesn't isolate pre-claim information and the headline 22% drop is a mislabeled micro-F1 number. read the letter →

arxiv 2507.03671 v1 pith:V2CQPE25 submitted 2025-07-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords factcheckinglargelanguagemodelsbenchmarkdatasetinformationleakageagenticpipelineclaimverificationpoliticalclaimsquestiondecomposition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that existing fact-checking benchmarks overstate model ability because their evidence passages contain the fact-checker's verdict and interpretive commentary. To match real-time verification, it introduces PFO, a five-class benchmark of 2,982 political claims whose evidence has been manually stripped of all post-claim analysis. On PFO, zero-shot LLMs lose about 22% macro-F1 on average compared with the unfiltered version, quantifying how much of their apparent skill came from annotator cues. The paper then proposes RAV, a three-agent pipeline that decomposes a claim into subquestions, answers them from evidence, and generates a label, and reports that it outperforms published baselines on RAWFC and HOVER while degrading less than baselines on PFO.

What carries the argument

The argument is carried by two constructed objects: PFO, a manually scrubbed five-class benchmark, and RAV, a three-agent agentic pipeline. PFO operationalizes 'no leakage' by removing post-claim verdict sentences and annotator commentary from PolitiFact evidence, leaving only facts that existed before the claim was published. RAV breaks a claim into a sequence of subquestions generated without access to evidence, answers each subquestion from the gold evidence, and then labels based on the question-answer history; its iterative question generation, mixing true/false verification questions with open inquiry questions, is what allows it to generalize across domains and label granularities.

What would settle it

Take a random sample of PFO claims, retrieve the actual documents, transcripts, or data releases that existed on or before the claim date, and compare zero-shot macro-F1 on those true pre-claim documents with macro-F1 on PFO's scrubbed evidence; a large gap would show that PFO's manual deletion does not reproduce pre-claim information conditions.

Watch

Extended reading notes

Core claim

The central claim is that leakage in fact-checking datasets materially inflates LLM evaluation: when evidence is reduced to factual content that existed before the claim was published, models perform far worse than when the evidence includes the fact-checker's post-publication analysis. PFO operationalizes this by manually deleting verdict sentences, label definitions, and interpretive commentary from PolitiFact articles, keeping only pre-claim facts. The paper's pipeline RAV—Recon, Answer, Verify—uses a question generator that iteratively asks subquestions without seeing the evidence, an answer generator that answers each subquestion from the evidence, and a label generator that predicts the final veracity label from the question-answer history. RAV outperforms state-of-the-art baselines on RAWFC and HOVER, and on FEVEROUS it trails ProgramFC; it also shows a smaller performance drop on PFO than the zero-shot baselines, with the best backbone losing only 7.36% macro-F1.

Load-bearing premise

The load-bearing premise is that manually deleting verdict-like sentences and annotator commentary from fact-checking articles leaves evidence equivalent to what a fact-checker would have had before the claim was published.

Editorial extensions

If this is right

  • Benchmark scores on unfiltered fact-checking datasets overstate real-world performance by roughly a fifth in macro-F1, because models exploit verdict cues that would not exist during real-time verification.
  • Fact-checking systems should be evaluated on evidence that predates the claim, and PFO provides one such test set for political claims with five label granularities.
  • An iterative question-answer decomposition improves veracity prediction across 2-class, 3-class, and 5-class settings, outperforming program-guided and hierarchical baselines on most tested benchmarks.
  • Allowing the question generator to mix verification and inquiry questions helps more than using either type alone, and iterative questioning beats generating all questions at once.
  • The number of subquestions RAV generates tracks dataset difficulty classes such as HOVER's hop count, suggesting a measurable notion of claim reasoning complexity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same manual scrubbing procedure could be applied to other fact-checking corpora, and PFO's measured 17.79% average length reduction offers a rough bound on how much leakage those datasets may contain.
  • Editorial inference: if PFO's filtering faithfully simulates pre-claim evidence, then the 22% gap is a direct measure of model reliance on leaked cues; a natural validation would compare PFO against true contemporaneous documents from the claim date.
  • Editorial inference: RAV's subquestion counts could serve as an automatic difficulty score, letting dataset curators stratify benchmarks by reasoning complexity without additional human annotation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper makes two main contributions. First, it introduces Politi-Fact-Only (PFO), a 5-class fact-checking dataset of 2,982 political claims derived from PolitiFact, in which the authors manually removed post-claim analysis and annotator cues. The stated aim is to evaluate models using only information available before the claim was verified, and the paper reports that zero-shot LLMs drop by 22% in macro-F1 on PFO relative to its unfiltered version. Second, it proposes RAV (Recon-Answer-Verify), an agentic pipeline with a question generator, an answer generator, and a label generator that iteratively decomposes claims into sub-questions; RAV is evaluated on RAWFC, LIAR-RAW, PFO, HOVER, and FEVEROUS, with reported improvements over published baselines on RAWFC and HOVER. The paper also includes ablations of question-generation strategies, question types, prompt sensitivity, and a human error analysis of RAV outputs.

Significance. If the PFO premise is accepted, the resource would be valuable: measuring the effect of annotator cues on LLM fact-checking is an important and understudied problem, and the manual curation effort, the inter-annotator agreement (Fleiss' kappa of 0.71), and the explicit documentation of annotation guidelines are strengths. The RAV pipeline is a reasonable agentic design, and the ablation study comparing iterative versus all-at-once question generation and verification versus inquiry questions is informative. The prompt-sensitivity analysis in Appendix A and the human error attribution in Appendix D are also useful practical contributions. However, the central quantitative claims, especially the 22% macro-F1 drop and the unconditional "outperforms state of the art" statements, are not fully supported by the tables as presented, and the PFO construction's core premise requires additional evidence before the headline results can be interpreted as measuring real-time verification.

major comments (3)
  1. [Section 4, Figures 11-12, Limitation] The claim that PFO contains "only the information that would have been available prior to the claim's verification" is not established. The evidence is extracted from fact-checking articles written after the claim, and even after manual deletion the retained text includes post-claim reasoning and verdict-relevant arithmetic. For example, Figure 11 keeps "Do the math: the sum exceeds $490.2 billion, much higher than even the highest estimate" and "That produces a total of $363 billion, well below the lowest estimate," and Figure 12 keeps "But this was no ordinary collision; it involved multiple vehicles that were not all pictured," which is an interpretive rebuttal of the claim. The Limitation section explicitly concedes that "some sentences could not be eliminated without compromising the context necessary to support or refute the claim." Therefore the 22% drop reported in the abstract is not demonstrably the cost of removing annotator cues; it may reflect partially scrubbed post-hoc articles. To make the PFO claim load-bearing, the authors should either re-scope the dataset description as "manually scrubbed PolitiFact articles" or provide a direct test of residual cue leakage, such as a blind annotator study measuring whether retained evidence still reveals the verdict, or a comparison against genuinely pre-claim corpora (e.g., WatClaimCheck or CofCED-style documents).
  2. [Abstract, Section 8, Table 2] The headline "average performance drop of 22% in terms of macro-f1" is not reproducible from Table 2. The macro-F1 drops for the three models are: Mistral-7B-v0.3, 0.37 to 0.26 (absolute 0.11); LLaMA-3.1-8b, 0.48 to 0.21 (absolute 0.27); Gemma-2-9b, 0.59 to 0.20 (absolute 0.39). The average absolute macro-F1 drop is approximately 0.257, and the relative drops are much larger (roughly 30%, 56%, and 66%). The value 0.22 matches the average absolute micro-F1 drop (0.45 to 0.34, 0.51 to 0.28, and 0.60 to 0.28). The paper should correct this claim, report both relative and absolute drops, and state explicitly which macro-F1 columns support the abstract's numbers. This is a load-bearing error because the 22% figure is prominently presented in the abstract and contribution list.
  3. [Section 8, Table 3, Table 4] The claim that RAV "outperforms state-of-the-art approaches on RAWFC by 25.28%" is only true for the phi-4 backbone: RAV(phi-4) on RAWFC is 0.6753 versus HiSS's 0.5390, a 25.3% relative improvement. The abstract does not specify the backbone, and Table 3 shows that RAV(LLaMA-3.1-70b) improves over HiSS by only about 9.8% on RAWFC. Similarly, on HOVER 2-hop, RAV(phi-4) at 0.7558 is slightly below ProgramFC's 0.7565, so the "1.54%" improvement is only from the 70B model. These statements need qualification by backbone. In addition, the baselines ProgramFC and HiSS use text-davinci-003 (175B), so the comparison conflates model choice with method effectiveness; the paper should acknowledge this confound and, ideally, report variance or significance across multiple runs, since all reported numbers appear to come from single executions.
minor comments (6)
  1. [Table 9] Table 9 reports a total of 2,981 instances for the unfiltered set, while Section 4 and Table 1 state 2,982; this inconsistency should be fixed.
  2. [Section 4] The sentence "On this contains around 21k instances" is grammatically broken and should be rewritten.
  3. [Table 3] The table header labels all columns "Macro-F1," but the ProgramFC baselines for HOVER are likely accuracy values from the original paper; the authors should clarify which metric is being reported for each baseline.
  4. [Section 5, Algorithm 1] The notation r*_QG and r*_LG is used in Algorithm 1 but not defined in the text; please define these symbols when the agents are introduced.
  5. [Appendix B.3] Instruction 5 says to write "yes" or "no" in the "leaked" field, but it is unclear whether "yes" indicates that the evidence required changes; please clarify the intended semantics.
  6. [Section 6] For RAWFC and LIAR-RAW, the authors follow prior work in treating author-written explanations as gold evidence; since those explanations may themselves contain verdict cues, the cross-dataset comparison is not apples-to-apples, and this should be stated as a limitation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: RAV is benchmarked against external published results, and PFO's filtered-versus-unfiltered contrast is an empirical comparison rather than a fitted prediction.

full rationale

The paper's central pipeline (RAV) is evaluated against RAWFC, HOVER, and FEVEROUS using published baseline numbers, so its claimed gains are externally anchored and not derived from its own outputs. The PFO dataset is a manual curation of PolitiFact articles; the reported 22% macro-F1 drop compares model performance on the authors' filtered data against the unfiltered source text. This is a measured effect of the curation, not a parameter fitted to a subset and then renamed a prediction. The paper has no self-citations and invokes no uniqueness theorem from its own prior work. The weaker point is that PFO's guarantee of 'only information available prior to verification' rests on the assumption that manual deletion removes all post-claim analysis; the retained 'Do the math' sentence in Figure 11 and the interpretive sentence in Figure 12, along with the Limitation's concession that 'some sentences could not be eliminated,' raise validity concerns about the 22% drop. But validity of the construct is a correctness/leakage issue, not an equation-level circularity: the paper does not define PFO's evidence in terms of the model scores, nor does it predict a number that is identical to its construction by definition. Because the main benchmark claims are externally anchored and no derivation step reduces to its own input, the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claims rest on subjective curation by three paid annotators (Fleiss kappa 0.7092 on 200 instances, 18 instances dropped for insufficient remaining context), on hyperparameters and a pipeline variant selected using the evaluation data itself, and on the unverified assumption that deleting sentences recovers pre-claim evidence. RAV's external benchmark results on RAWFC, HOVER, and FEVEROUS are the most independent part of the paper. No invented entities are introduced; the three agents are software components.

free parameters (3)
  • Maximum QA iterations (k) = 10
    Set from Figure 3, which plots macro-F1 against k on the evaluation datasets themselves; the text says performance plateaus or degrades beyond 10. Tuning a hyperparameter on the benchmarks used for final reporting inflates results.
  • Agentic variant (P2, T1 and T2) = chosen variant
    The final pipeline was selected from Table 4, an ablation run on the test sets of LIAR-RAW, RAWFC, FEVEROUS, and HOVER. Choosing the best test-set performer as 'ours' is selection on the evaluation data.
  • Zero-shot prompt (P6) = P6
    Best of 7 prompt variants in Table 6, selected on the validation set of the unfiltered PFO. Prompt phrasing alone shifts F1 by up to 10 points, so the choice matters for every zero-shot number in the paper.
assumptions (4)
  • domain assumption PolitiFact verdicts are treated as ground truth for claim veracity.
    Section 2 defines the prediction task against gold labels; Section 4 builds PFO directly from Misra (2022) labels. PolitiFact ratings are editorial judgments, a standard but unexamined assumption.
  • ad hoc to paper Deleting verdict-like sentences leaves evidence equivalent to pre-claim information.
    Stated in Section 4 ('We retain only the facts that existed before the claim was published') and implemented through Appendix B.3 rules. The filtered examples in Figures 10-13 and the Limitation show residual interpretive content, so the equivalence does not hold strictly.
  • ad hoc to paper Token-length reduction is a proxy for leaked content removed.
    Table 1's caption: 'Content length reduction confirms that 17.79% of the original data consisted of commentary or verdict cues.' Removing 17.79% of tokens does not establish that exactly those tokens were cues.
  • domain assumption Gold evidence in RAWFC and LIAR-RAW, as author-written explanations, is usable evidence.
    Section 6 follows Yang et al. (2022) in using author-written explanations as gold evidence. Those explanations may themselves contain verdict-adjacent language, the same leakage the paper warns about.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Recon, Answer, Verify: Agents in Search of Truth." pith.science (2026). https://pith.science/paper/V2CQPE25

@misc{pith2026250703671,
  author       = {Pith},
  title        = {Pith review of: Recon, Answer, Verify: Agents in Search of Truth},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2CQPE25}},
  note         = {Machine review of arXiv:2507.03671}
}
read the original abstract

Automated fact checking with large language models (LLMs) offers a scalable alternative to manual verification. Evaluating fact checking is challenging as existing benchmark datasets often include post claim analysis and annotator cues, which are absent in real world scenarios where claims are fact checked immediately after being made. This limits the realism of current evaluations. We present Politi Fact Only (PFO), a 5 class benchmark dataset of 2,982 political claims from politifact.com, where all post claim analysis and annotator cues have been removed manually. This ensures that models are evaluated using only the information that would have been available prior to the claim's verification. Evaluating LLMs on PFO, we see an average performance drop of 22% in terms of macro f1 compared to PFO's unfiltered version. Based on the identified challenges of the existing LLM based fact checking system, we propose RAV (Recon Answer Verify), an agentic framework with three agents: question generator, answer generator, and label generator. Our pipeline iteratively generates and answers sub questions to verify different aspects of the claim before finally generating the label. RAV generalizes across domains and label granularities, and it outperforms state of the art approaches on well known baselines RAWFC (fact checking, 3 class) by 25.28%, and on HOVER (encyclopedia, 2 class) by 1.54% on 2 hop, 4.94% on 3 hop, and 1.78% on 4 hop, sub categories respectively. RAV shows the least performance drop compared to baselines of 16.3% in macro f1 when we compare PFO with its unfiltered version.

Figures

Figures reproduced from arXiv: 2507.03671 by the authors.

Figure 1
Figure 1. An instance from the PFO dataset where sen [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. This diagram illustrates the workflow of the RAV pipeline for an instance (claim & evidence pair), [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Macro F1 score of the RAV pipeline across [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Distribution of number of sub-questions per In [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: This figure presents the fault attribution across different datasets and agent types involved in the claim [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: QGagent Prompt used in our RAV pipeline, we give the clean instructions to generate the follow-up question at each iteration, stop the iteration by outputting stop_iteration. We also provided 8 in-context examples in the prompt [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: AGagent Prompt used in our RAV pipeline, we give the clean instructions to generate the answer in 10 words, we also instruct to look for indirect answers present in the context. We also instruct AGagent to completely rely on the evidence [PITH_FULL_IMAGE:figures/full_…
Figure 8
Figure 8. Figure 8: AGagent Prompt used in our RAV pipeline, we give a short description of the task and provide 8 in-context examples in the prompt, we instruct AGagent to first generate a reasoning to connect the claim and generated questions and answers and then predict the label [PIT…
Figure 9
Figure 9. Figure 9: A true instance from the PFO dataset. Label: mostly-true Claim: The failings in our civil service are encouraged by a system that makes it very difficult to fire someone even for gross misconduct. Evidence: Sen. John McCain, the arizona republican, overstates the probl…
Figure 10
Figure 10. Figure 10: A mostly true instance from the PFO dataset. [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: A half true instance from the PFO dataset. [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: A mostly false instance from the PFO dataset. [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]
Figure 13
Figure 13. Figure 13: A false instance from the PFO dataset [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

    cs.CL 2025-11 unverdicted novelty 5.0 of 10

    REFLEX improves explainable fact-checking by using verdict-anchored style control and self-disagreement signals to disentangle fact from style in LLM outputs, achieving SOTA results with minimal self-refined samples.

Reference graph

Works this paper leans on

21 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Tariq Alhindi, Savvas Petridis, and Smaranda Muresan. 2018. https://doi.org/10.18653/v1/W18-5513 Where is your evidence: Improving fact-checking by justification modeling . In Proceedings of the First Workshop on Fact Extraction and VER ification ( FEVER ) , pages 85--90, Brussels, Belgium. Association for Computational Linguistics

  4. [4]

    Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. https://doi.org/10.18653/v1/D19-1475 M ulti FC : A real-world multi-domain dataset for evidence-based fact checking of claims . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing an...

  5. [5]

    Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. 2023. Lm vs lm: Detecting factual errors via cross examination. arXiv preprint arXiv:2305.13281

  6. [6]

    Lucas Graves. 2018. Understanding the promise and limits of automated fact-checking. Reuters Institute for the Study of Journalism

  7. [7]

    Ashim Gupta and Vivek Srikumar. 2021. https://doi.org/10.18653/v1/2021.acl-short.86 X -fact: A new benchmark dataset for multilingual fact checking . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 675--682,...

  8. [8]

    Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.309 H o V er: A dataset for many-hop fact extraction and claim verification . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441--3460, Online. Association for Computational Linguistics

Show all 21 references
  1. [9]

    Kashif Khan, Ruizhe Wang, and Pascal Poupart. 2022. https://doi.org/10.18653/v1/2022.acl-long.92 W at C laim C heck: A new dataset for claim entailment and inference . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...

  2. [10]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2024. Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA. Curran Associates Inc

  3. [11]

    Justin Matthew Wren Lewis, Andy Williams, Robert Arthur Franklin, James Thomas, and Nicholas Alexander Mosdell. 2008. The quality and independence of british journalism

  4. [12]

    Rishabh Misra. 2022. https://doi.org/10.13140/RG.2.2.29923.22566 Politifact fact check dataset

  5. [13]

    Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, and Preslav Nakov. 2023. https://arxiv.org/abs/2305.12744 Fact-checking complex claims with program-guided reasoning

  6. [14]

    Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. 2017. https://doi.org/10.18653/v1/D17-1317 Truth of varying shades: Analyzing language in fake news and political fact-checking . In Proceedings of the 2017 Conference on Empirical Methods in Natural ...

  7. [15]

    Daniel Russo, Serra Sinem Tekiroğlu, and Marco Guerini. 2023. https://doi.org/10.1162/tacl_a_00601 Benchmarking the Generation of Fact Checking Explanations . Transactions of the Association for Computational Linguistics, 11:1250--1264

  8. [16]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. https://doi.org/10.18653/v1/N18-1074 FEVER : a large-scale dataset for fact extraction and VER ification . In Proceedings of the 2018 Conference of the North A merican Chapter of the Associatio...

  9. [17]

    Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science, 359(6380):1146--1151

  10. [18]

    William Yang Wang. 2017. https://doi.org/10.18653/v1/P17-2067 `` liar, liar pants on fire '' : A new benchmark dataset for fake news detection . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 422--426,...

  11. [19]

    Zhiwei Yang, Jing Ma, Hechang Chen, Hongzhan Lin, Ziyang Luo, and Chang Yi. 2022. https://aclanthology.org/2022.coling-1.230 A coarse-to-fine cascaded evidence-distillation neural network for explainable fake news detection . In Proceedings of the 29th International Conference...

  12. [20]

    Yirong Zeng, Xiao Ding, Yi Zhao, Xiangyu Li, Jie Zhang, Chao Yao, Ting Liu, and Bing Qin. 2024. https://aclanthology.org/2024.lrec-main.1239/ RU 22 F act: Optimizing evidence for multilingual explainable fact-checking on R ussia- U kraine conflict . In Proceedings of the 2024 ...

  13. [21]

    Xuan Zhang and Wei Gao. 2023. http://arxiv.org/abs/2310.00305 Towards llm-based fact verification on news claims with a hierarchical step-by-step prompting method

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.