Pith. sign in

REVIEW 3 major objections 4 minor 66 references

Fact-checking LLMs show a confidence paradox: smaller models overconfident, larger models overcautious.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 20:05 UTC pith:2YCQSZRS

load-bearing objection The paper's empirical core—small models answer more often but less accurately, large models abstain more but are more accurate—is real and worth engaging, but the Dunning-Kruger framing overreaches because certainty rate is not confidence. the 3 major comments →

arxiv 2509.08803 v1 pith:2YCQSZRS submitted 2025-09-10 cs.SI cs.AIcs.CLcs.CY

Scaling Truth: The Confidence Paradox in AI Fact-Checking

classification cs.SI cs.AIcs.CLcs.CY
keywords fact-checkinglarge language modelsDunning-Kruger effectselective classificationcertainty ratemultilingual evaluationmisinformationmodel confidence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper evaluates nine large language models on 5,000 real-world claims vetted by professional fact-checkers across 47 languages, using prompt styles that mimic ordinary users. It finds a systematic trade-off: smaller models give definite verdicts for most claims but are correct only about 60% of the time, while larger models are more accurate (up to about 89%) but abstain on most claims. The authors call this a Dunning-Kruger-like 'confidence-competence paradox' and argue it matters because resource-constrained fact-checkers tend to rely on the smaller, overconfident models. They also show performance drops for non-English claims and claims from the Global South, and that a reasoning-focused model can achieve both high accuracy and high certainty, though at much higher cost.

Core claim

Across nine models and four prompting strategies, the authors report an inverse relationship between selective accuracy—the fraction of definitive True/False verdicts that are correct—and certainty rate—the fraction of claims for which the model gives a definitive verdict. Smaller models commit to verdicts on up to 88% of claims but land around 60% selective accuracy; larger models reach up to 89% selective accuracy but issue definitive verdicts on fewer than 40% of claims. The pattern persists when only common claims are analyzed, when MMLU scores are used as the ability measure, and on claims published after training cutoffs. The authors interpret this as a statistical analogue of the Dunn

What carries the argument

The core machinery is the three-metric evaluation framework: selective accuracy (correctness among definitive verdicts), abstention-friendly accuracy (correct verdicts plus appropriate abstentions), and certainty rate (the proportion of claims answered definitively, also called coverage in selective classification). The load-bearing pattern is the inverse correlation between certainty rate and selective accuracy across model sizes, visualized in quartile analyses that place models on an 'actual ability vs. perceived ability' plane analogous to Dunning-Kruger plots.

Load-bearing premise

The paper treats the rate of definitive answers as a measure of a model's perceived confidence; if abstaining reflects safety training, prompt formatting, or other factors rather than genuine uncertainty, the confidence-competence paradox may be an artifact of the metric.

What would settle it

Ask the same models to answer every claim with just 'True' or 'False' (no abstention option) and compare selective accuracy. If a large model's selective accuracy stays near 89% when forced to answer everything, then its low certainty rate is a response-format choice rather than a true competence gap; conversely, if a small model's selective accuracy drops further when it is forced to commit on claims it would otherwise abstain on, the overconfidence pattern is real and even stronger.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Resource-constrained organizations that rely on small, cheap models face a higher risk of confident misinformation propagation.
  • Structured system prompts improve selective accuracy for all models but make smaller models overcommit, whereas larger models become more accurate without becoming more certain.
  • Non-English and Global South claims show lower selective accuracy and certainty for most models, threatening to widen information inequalities.
  • A reasoning-based model achieves both high accuracy and high certainty, suggesting architectural improvements can mitigate the paradox, but at a cost of $88.75 per 1,000 claims.
  • The findings support policy attention to equitable access to reliable fact-checking tools, e.g., under the EU AI Act.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The Dunning-Kruger analogy is a statistical mapping of certainty rate to perceived ability; it does not establish that models experience subjective confidence, and the same pattern could arise from different abstention policies across model families.
  • A natural testable extension is to elicit self-reported confidence (e.g., verbalized probabilities) from the same models and check whether it tracks certainty rate; if it does not, the paradox may be a metric artifact.
  • The regional and linguistic disparities may partly reflect training-data skew, and could be probed by testing the same models on claims from the Global South after fine-tuning on balanced data.
  • The cost-performance trade-off suggests that subsidized access to larger models, or distillation of reasoning models into smaller ones, could be a more direct remedy than calibration alone.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a large-scale evaluation of nine LLMs (Llama-2 family, Mistral/Mixtral, GPT-3.5/4/4o, o1-preview) on 5,000 real-world claims verified by professional fact-checking organizations across 47 languages. Using four prompting strategies (three user-level prompts in the claim's language and one IFCN-style system prompt with structured output), the authors annotate each model response as True/False/Other and compute selective accuracy, abstention-friendly accuracy, and certainty rate. The central claim is a 'confidence-competence paradox' reminiscent of the Dunning-Kruger effect: smaller models show high certainty rates but lower selective accuracy, while larger models show lower certainty but higher accuracy. The paper also documents performance degradation on non-English and Global South claims and discusses cost, equity, and policy implications. The main empirical contribution is the multilingual benchmark and the demonstration that prompt format and model scale affect the accuracy-abstention trade-off.

Significance. If the central interpretation holds, the finding is practically important: it implies that the cheapest, most accessible LLM fact-checking tools are also the most overconfident, creating a systemic risk for resource-constrained organizations. The study's strengths are real: a multilingual claim set grounded in professional fact-checker verdicts, human annotation with reported inter-annotator agreement, multiple prompting strategies, post-training-cutoff robustness checks, and statistical corrections for multiple comparisons. The paper also makes useful methodological moves by distinguishing selective accuracy from abstention-friendly accuracy. However, the headline contribution depends on identifying certainty rate with the model's 'perceived confidence.' That load-bearing assumption is not established by the evidence in the manuscript, because the 'Other' category conflates epistemic uncertainty with inability to verify, refusal, non-responsiveness, and hesitation. The observed inverse pattern may therefore reflect differences in abstention style rather than a confidence-competence trade-off. The Dunning-Kruger analogy is also applied at the aggregate level across eight models, no

major comments (3)
  1. [§2.2, §4.2, §A.2.1] The paper's central claim—that certainty rate measures an LLM's 'self-perceived confidence' (Fig. 3; §2.3)—is not supported by the manuscript's own operationalization. Certainty rate is defined in §4.2 as the proportion of responses not labeled 'Other,' and 'Other' is defined in §A.2.1 to include five distinct subcategories: Ambiguity/Uncertainty, Partial Assessment, Inability to Verify, Non-Responsive/Irrelevant, and Hesitation/Avoidance. A model may abstain because it lacks real-time information, because safety training makes it avoid a verdict, or because its output format does not match the annotation scheme; none of these is equivalent to low epistemic confidence. The paper itself notes in §2.2 that prompt format changes certainty rates and that smaller models 'over-interpret structured directives,' which shows that the metric tracks instruction-following and response style, not jus
  2. [§2.3, Fig. 3] Even if certainty rate were a valid confidence measure, the Dunning-Kruger framing is stronger than the evidence. The Dunning-Kruger effect is a within-person relation between actual performance and self-assessed ability, typically demonstrated across quartiles or individuals facing the same task. Here the evidence is an aggregate inverse correlation across eight model families (Fig. 3, top-left), averaging over prompts and languages. The robustness check using a common claim set (Fig. B3) uses small samples (n=357, 333, 526, 846) and still aggregates at the model level. The paper should either reframe the result as a 'scale-dependent coverage–accuracy trade-off' or add an analysis at the level at which Dunning-Kruger is normally tested—for example, by eliciting per-claim confidence estimates from the same model and comparing them to that model's correctness on those claims. Without such
  3. [§2.4, Table C4] The Global North/South selective-accuracy comparisons use unequal and sometimes small subsets because 'Other' responses are excluded. For example, GPT-4 has nNorth=482 and nSouth=205, and GPT-4o has nNorth=509 and nSouth=224 (Table C4). Selective accuracy is computed on different claim sets for the two regions, so a model may appear to degrade in the South simply because the subset of claims on which it committed is harder, not because of a systematic regional deficit. The paper reports tests on these disparate subsets without a common-claims analysis. Since the abstract and discussion foreground regional disparities (e.g., 'most pronounced for non-English languages and claims originating from the Global South'), the authors should report the overlap of committed claims between regions or model the abstention and correctness decisions jointly (e.g., a two-stage analysis).
minor comments (4)
  1. [§2.4 vs. §A.3] The inter-annotator agreement for the Global North/South regional classification is reported as Cohen's Kappa = 0.802 in §2.4 but as 0.84 in §A.3. The discrepancy should be reconciled.
  2. [Abstract / §2.1] The abstract says 'over 240,000 human annotations.' With eight models × three prompts × 5,000 claims plus the o1-preview Prompt 4 condition, and two annotators per response, the implied total should be computed and stated explicitly. Please clarify the counting, or the number will appear inconsistent.
  3. [§A.5] The post-training robustness check is conducted exclusively with Prompt 4 and with claims published after October 2023. This is a reasonable check, but the discussion in §A.5 should note the smaller scope and that prompt-specific effects may differ for user-level prompts. Also, 'postdating the knowledge cutoff of the latest model' means the claims are post-cutoff for all models, which is fine, but the phrase 'postdating training cutoffs' in the abstract should be qualified.
  4. [Supplementary Fig. B5 and §4.2] The supplementary figures use 'abstention-free accuracy' while the main text and the metric definition in §4.2 use 'abstention-friendly accuracy.' Please standardize the terminology.

Circularity Check

0 steps flagged

No significant circularity: the paper is an empirical evaluation with transparently defined metrics; the Dunning-Kruger framing is an explicitly conditional analogy, not a derivation.

full rationale

Walked the claimed derivation chain: dataset construction (Sec. 4.1), response annotation (Sec. 2.1, Sec. A.2), metric definitions (Sec. 4.2), cross-model/cross-prompt comparisons (Sec. 2.2–2.4), and the Dunning-Kruger interpretation (Sec. 2.3). At no point is a reported quantity derived from another quantity it was defined to produce. Selective accuracy, abstention-friendly accuracy, and certainty rate are each defined directly from the annotated response contingency table in Sec. 4.2; certainty rate is explicitly identified with the standard selective-classification notion of coverage and is only called a proxy for confidence in Sec. 2.2. The inverse accuracy-certainty pattern is an empirical correlation across models and prompts (Fig. 3), and the paper supports it with robustness checks on common claim subsets, MMLU scores, and post-training claims. The phrase 'resembles the Dunning-Kruger effect' is presented as an analogy, with the sentence explicitly conditional on the chosen measures ('when selective accuracy and certainty rates are used as measures of actual and perceived ability, respectively'). There are no fitted parameters later relabeled as predictions, no load-bearing self-citations (the reference list contains no author self-citations), and no uniqueness or ansatz smuggled in via citation. The only substantive concern—whether certainty rate validly operationalizes perceived confidence—is a construct-validity limitation of the proxy, not circularity, because the proxy is disclosed and the observed relationship is not an identity. The paper is therefore self-contained as an empirical study rather than a circular derivation.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The paper introduces no new mathematical entities or free parameters. It relies on empirical assumptions about data provenance, translation fidelity, and annotation reliability. The main conceptual entity is the operationalization of 'certainty rate' as perceived confidence, which is a definitional choice rather than an invented physical or formal object.

axioms (3)
  • domain assumption Ground-truth labels from Google Fact Check and fact-checking organizations are correct for the claims used.
    All accuracy metrics compare model outputs against these labels; no independent verification of the labels is performed. See Section 2.1.
  • domain assumption Machine translation via Google Translate preserves the semantics of model responses for annotation purposes.
    Non-English model responses were translated to English before human annotation (Section A.1 and Section 2.1). Translation errors could change the classification of a response.
  • domain assumption Human annotations of model responses into True/False/Other are reliable despite the subjective line between definitive and uncertain answers.
    Cohen's Kappa of 0.84 is substantial but not perfect, and discrepancies were resolved by consensus, which can introduce systematic bias. See Section 2.1 and A.2.

pith-pipeline@v1.3.0-alltime-deepseek · 32083 in / 6345 out tokens · 71373 ms · 2026-08-04T20:05:12.739167+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Scaling Truth: The Confidence Paradox in AI Fact-Checking." pith.science (2026). https://pith.science/paper/2YCQSZRS

@misc{pith2026250908803,
  author       = {Pith},
  title        = {Pith review of: Scaling Truth: The Confidence Paradox in AI Fact-Checking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2YCQSZRS}},
  note         = {Machine review of arXiv:2509.08803}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet their effectiveness across global contexts remains uncertain. We systematically evaluate nine established LLMs across multiple categories (open/closed-source, multiple sizes, diverse architectures, reasoning-based) using 5,000 claims previously assessed by 174 professional fact-checking organizations across 47 languages. Our methodology tests model generalizability on claims postdating training cutoffs and four prompting strategies mirroring both citizen and professional fact-checker interactions, with over 240,000 human annotations as ground truth. Findings reveal a concerning pattern resembling the Dunning-Kruger effect: smaller, accessible models show high confidence despite lower accuracy, while larger models demonstrate higher accuracy but lower confidence. This risks systemic bias in information verification, as resource-constrained organizations typically use smaller models. Performance gaps are most pronounced for non-English languages and claims originating from the Global South, threatening to widen existing information inequalities. These results establish a multilingual benchmark for future research and provide an evidence base for policy aimed at ensuring equitable access to trustworthy, AI-assisted fact-checking.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

66 extracted references · 3 canonical work pages

  1. [1]

    Nature Human Behavior 8(9), 1643–1655 (2024) https://doi.org/10.1038/s41562-024-01959-9

    Burton, J.W., Lopez-Lopez, E., Hechtlinger, S., Rahwan, Z., Aeschbach, S., Bakker, M.A., Becker, J.A., Berditchevskaia, A., Berger, J., Brinkmann, L., Flek, L., Herzog, S.M., Huang, S., Kapoor, S., Narayanan, A., Nussberger, A.M., Yasseri, T., Nickl, P., Almaatouq, A., Hahn, U., Kurvers, R.H.J.M., Leavy, S., Rahwan, I., Siddarth, D., Siu, A., Woolley, A.W...

  2. [2]

    Journalism0(0) (2025) https://doi.org/10.1177/14648849251317150

    Jones, B., Jones, R.: Action research at the bbc: Interrogating artificial intel- ligence with journalists to generate actionable insights for the newsroom. Journalism0(0) (2025) https://doi.org/10.1177/14648849251317150

  3. [3]

    Proc Natl Acad Sci U S A121(50), 2322823121 (2024) https://doi.org/10.1073/pnas.2322823121

    DeVerna, M.R., Yan, H.Y., Yang, K.C., Menczer, F.: Fact-checking information from large language models can decrease headline discernment. Proc Natl Acad Sci U S A121(50), 2322823121 (2024) https://doi.org/10.1073/pnas.2322823121 . Epub 2024 Dec 4

  4. [4]

    Technical Report 24-93, Swiss Finance Institute (February 2024)

    Leippold, M., Vaghefi, S., Muccione, V., Bingler, J., Stammbach, D., Cole- santi Senni, C., Ni, J., Wekhof, T., Yu, T., Schimanski, T., Gostlow, G., Luterbacher, J., Huggel, C.: Automated fact-checking of climate claims with large language models. Technical Report 24-93, Swiss Finance Institute (February 2024). Forthcoming, Nature Climate Action. https://...

  5. [5]

    Proceedings of the National Academy of Sci- ences120(30), 2305016120 (2023) https://doi.org/10.1073/pnas.2305016120 https://www.pnas.org/doi/pdf/10.1073/pnas.2305016120

    Gilardi, F., Alizadeh, M., Kubli, M.: Chatgpt outperforms crowd work- ers for text-annotation tasks. Proceedings of the National Academy of Sci- ences120(30), 2305016120 (2023) https://doi.org/10.1073/pnas.2305016120 https://www.pnas.org/doi/pdf/10.1073/pnas.2305016120

  6. [6]

    AI & Society (2025) https://doi.org/10.1007/ s00146-025-02199-9

    Vetter, M.A., Jiang, J., McDowell, Z.J.: An endangered species: how llms threaten wikipedia’s sustainability. AI & Society (2025) https://doi.org/10.1007/ s00146-025-02199-9

  7. [7]

    https://arxiv.org/abs/2005.14165

    Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Nee- lakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Win- ter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford...

  8. [8]

    https://arxiv.org/abs/2307.09288

    Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bash- lykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C.C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabs...

  9. [9]

    Accessed: 2024-11-15 (2024)

    Reid, L.: Generative AI in Search: Let Google do the searching for you. Accessed: 2024-11-15 (2024). https://blog.google/products/search/ generative-ai-google-search-may-2024/

  10. [10]

    Accessed: 2024- 12-07 (2024)

    Institute, T.P.: The State of Fact-Checkers 2023. Accessed: 2024- 12-07 (2024). https://www.poynter.org/wp-content/uploads/2024/04/ State-of-Fact-Checkers-2023.pdf

  11. [11]

    Accessed: 2024-11-12 (2024)

    Journalism, R.I.: Digital News Report 2024. Accessed: 2024-11-12 (2024). https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2024-06/ RISJ DNR 2024 Digital v10%20lr.pdf

  12. [12]

    Nature637(8047), 778–780 (2025) https://doi.org/10.1038/ d41586-025-00068-5

    Jones, N.: Ai hallucinations can’t be stopped - but these techniques can limit their damage. Nature637(8047), 778–780 (2025) https://doi.org/10.1038/ d41586-025-00068-5

  13. [13]

    Scientific Reports14, 16375 (2024) https://doi.org/10.1038/s41598-024-66708-4

    Hao, G., Wu, J., Pan, Q.,et al.: Quantifying the uncertainty of llm hallucina- tion spreading in complex adaptive social networks. Scientific Reports14, 16375 (2024) https://doi.org/10.1038/s41598-024-66708-4

  14. [14]

    ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730

    Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730

  15. [15]

    Nature Machine Intelligence7(2), 221–231 (2025) https://doi.org/10.1038/ s42256-024-00976-7

    Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L.W., Smyth, P.: What large language models know and what people think they know. Nature Machine Intelligence7(2), 221–231 (2025) https://doi.org/10.1038/ s42256-024-00976-7

  16. [16]

    https://arxiv.org/abs/2412.01459

    Brauner, P., Glawe, F., Liehner, G.L., Vervier, L., Ziefle, M.: Misalignments in AI Perception: Quantitative Findings and Visual Mapping of How Experts and the Public Differ in Expectations and Risks, Benefits, and Value Judgments (2024). https://arxiv.org/abs/2412.01459

  17. [17]

    Forbes (2025)

    Daniel, L.: The irony—ai expert’s testimony collapses over fake ai citations. Forbes (2025). Accessed: 2025-02-09

  18. [18]

    https://arxiv.org/abs/2404.09356 19

    Sathish, V., Lin, H., Kamath, A.K., Nyayachavadi, A.: LLeMpower: Understand- ing Disparities in the Control and Access of Large Language Models (2024). https://arxiv.org/abs/2404.09356 19

  19. [19]

    Accessed: 2025-02-15

    Whiting, K.: What is a small language model and how can businesses leverage this ai tool? World Economic Forum (2025). Accessed: 2025-02-15

  20. [20]

    In: Proceedings of the 30th International Conference on Intelligent User Interfaces

    Dudy, S., Tholeti, T., Ramachandranpillai, R., Ali, M., Li, T.J.-J., Baeza-Yates, R.: Unequal opportunities: Examining the bias in geographical recommendations by large language models. In: Proceedings of the 30th International Conference on Intelligent User Interfaces. IUI ’25, pp. 1499–1516. Association for Comput- ing Machinery, New York, NY, USA (2025...

  21. [21]

    Sheng, E., Chang, K.-W., Natarajan, P., Peng, N.: The woman worked as a babysitter: On biases in language generation. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJCNLP), pp. 3407–3412. Association for Computational Linguistics,...

  22. [22]

    Accessed: 2025-04-09 (2025)

    Meta Newsroom: Testing Begins for Community Notes on Facebook, Instagram and Threads. Accessed: 2025-04-09 (2025). https://about.fb.com/news/2025/03/ testing-begins-community-notes-facebook-instagram-threads/

  23. [23]

    Scientific Reports14, 26133 (2024) https://doi.org/10.1038/s41598-024-76900-1

    Porter, B., Machery, E.: Ai-generated poetry is indistinguishable from human- written poetry and is rated more favorably. Scientific Reports14, 26133 (2024) https://doi.org/10.1038/s41598-024-76900-1

  24. [24]

    Proceedings of the National Academy of Sciences120(11) (2023) https://doi.org/10.1073/pnas.2208839120 2208839120

    Jakesch, M., Hancock, J.T., Naaman, M.: Human heuristics for ai-generated lan- guage are flawed. Proceedings of the National Academy of Sciences120(11) (2023) https://doi.org/10.1073/pnas.2208839120 2208839120

  25. [25]

    Science Advances9(26), 1850 (2023) https://doi.org/10

    Spitale, G., Biller-Andorno, N., Germani, F.: Ai model gpt-3 (dis)informs us better than humans. Science Advances9(26), 1850 (2023) https://doi.org/10. 1126/sciadv.adh1850 https://www.science.org/doi/pdf/10.1126/sciadv.adh1850

  26. [26]

    Accessed: 2025-01-05 (2024)

    European Commission: AI Act. Accessed: 2025-01-05 (2024). https:// digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai

  27. [27]

    PsyArXiv (2023)

    Hoes, E., Altay, S., Bermeo, J.: Leveraging ChatGPT for Efficient Fact- Checking. PsyArXiv (2023). https://doi.org/10.31234/osf.io/qnjkf . osf.io/ preprints/psyarxiv/qnjkf v1

  28. [28]

    Frontiers in Artificial Intelligence7(2024) https://doi.org/10.3389/frai

    Quelle, D., Bovet, A.: The perils and promises of fact-checking with large language models. Frontiers in Artificial Intelligence7(2024) https://doi.org/10.3389/frai. 2024.1341697

  29. [29]

    https: //arxiv.org/abs/2411.10541 20

    He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F.X., Hasan, S.: Does Prompt Formatting Have Any Impact on LLM Performance? (2024). https: //arxiv.org/abs/2411.10541 20

  30. [30]

    https://arxiv.org/abs/2408.02442

    Tam, Z.R., Wu, C.-K., Tsai, Y.-L., Lin, C.-Y., Lee, H.-y., Chen, Y.-N.: Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models (2024). https://arxiv.org/abs/2408.02442

  31. [31]

    In: Ku, L.-W., Martins, A., Srikumar, V

    Ma, H., Xu, W., Wei, Y., Chen, L., Wang, L., Liu, Q., Wu, S., Wang, L.: EX- FEVER: A dataset for multi-hop explainable fact verification. In: Ku, L.-W., Martins, A., Srikumar, V. (eds.) Findings of the Association for Computational Linguistics: ACL 2024, pp. 9340–9353. Association for Computational Linguistics, Bangkok, Thailand (2024). https://doi.org/10...

  32. [32]

    In: Bouamor, H., Pino, J., Bali, K

    Pelrine, K., Imouza, A., Thibault, C., Reksoprodjo, M., Gupta, C., Christoph, J., Godbout, J.-F., Rabbany, R.: Towards reliable misinformation mitigation: Generalization, uncertainty, and GPT-4. In: Bouamor, H., Pino, J., Bali, K. (eds.) Proceedings of the 2023 Conference on Empirical Methods in Natu- ral Language Processing, pp. 6399–6429. Association fo...

  33. [33]

    In: 2023 IEEE International Conference on Data Mining Workshops (ICDMW), pp

    Li, Y., Zhai, C.: An Exploration of Large Language Models for Verifica- tion of News Headlines . In: 2023 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 197–206. IEEE Computer Society, Los Alamitos, CA, USA (2023). https://doi.org/10.1109/ICDMW60847.2023.00032 . https://doi.ieeecomputersociety.org/10.1109/ICDMW60847.2023.00032

  34. [34]

    DA WN.COM (2025)

    Dawn: Hey chatbot, is this true? ai ‘factchecks’ sow misinformation. DA WN.COM (2025)

  35. [35]

    In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval

    Trippas, J.R., Al Lawati, S.F.D., Mackenzie, J., Gallagher, L.: What do users really ask large language models? an initial log analysis of google bard interactions in the wild. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR ’24, pp. 2703–

  36. [36]

    https://www.semrush.com/blog/ chatgpt-search-insights/

    Kelly, B., Harsel, L.: Investigating ChatGPT Search: Insights from 80 Million Clickstream Records (2025). https://www.semrush.com/blog/ chatgpt-search-insights/

  37. [37]

    Accessed: 2024-11-03 (2024)

    International Fact-Checking Network: International Fact-Checking Network. Accessed: 2024-11-03 (2024). https://www.poynter.org/ifcn/

  38. [38]

    Journal of Personality and Social Psychology77(6), 1121–1134 (1999) https://doi.org/10

    Kruger, J., Dunning, D.: Unskilled and unaware of it: How difficulties in rec- ognizing one’s own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology77(6), 1121–1134 (1999) https://doi.org/10. 1037/0022-3514.77.6.1121 21

  39. [39]

    Accessed: 2024-10-02

    Google: Fact Check Tools - About. Accessed: 2024-10-02. https://toolbox.google. com/factcheck/about

  40. [40]

    https://arxiv.org/abs/2310.06825

    Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L.R., Lachaux, M.- A., Stock, P., Scao, T.L., Lavril, T., Wang, T., Lacroix, T., Sayed, W.E.: Mistral 7B (2023). https://arxiv.org/abs/2310.06825

  41. [41]

    https://arxiv.org/abs/2401.04088

    Jiang, A.Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D.S., Casas, D., Hanna, E.B., Bressand, F., Lengyel, G., Bour, G., Lam- ple, G., Lavaud, L.R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T.L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., Sayed, W.E.: Mixtral of Experts (2...

  42. [42]

    Accessed: 2025-02-09 (2025)

    OpenAI: OpenAI Models Documentation. Accessed: 2025-02-09 (2025). https: //platform.openai.com/docs/models

  43. [43]

    IEEE Trans

    Chow, C.K.: On optimum recognition error and reject tradeoff. IEEE Trans. Inf. Theor.16(1), 41–46 (1970) https://doi.org/10.1109/TIT.1970.1054406

  44. [44]

    https://doi.org/10.1109/TEC.1957.5222035

    Chow, C.K.: An optimum character recognition system using decision functions (1957). https://doi.org/10.1109/TEC.1957.5222035

  45. [45]

    https://arxiv.org/abs/2407.01032

    Traub, J., Bungert, T.J., L¨ uth, C.T., Baumgartner, M., Maier-Hein, K.H., Maier- Hein, L., Jaeger, P.F.: Overcoming Common Flaws in the Evaluation of Selective Classification Systems (2024). https://arxiv.org/abs/2407.01032

  46. [46]

    https://arxiv.org/abs/2407.16221

    Madhusudhan, N., Madhusudhan, S.T., Yadav, V., Hashemi, M.: Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models (2024). https://arxiv.org/abs/2407.16221

  47. [47]

    F AccT ’22, pp

    Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L.A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., Gabriel, I.: Taxonomy of risks posed by language models. F AccT ’22, pp. 214–22...

  48. [48]

    https://arxiv.org/abs/2109.07958

    Lin, S., Hilton, J., Evans, O.: TruthfulQA: Measuring How Models Mimic Human Falsehoods (2022). https://arxiv.org/abs/2109.07958

  49. [49]

    https://arxiv.org/abs/2405.04760

    Xu, H., Wang, S., Li, N., Wang, K., Zhao, Y., Chen, K., Yu, T., Liu, Y., Wang, H.: Large Language Models for Cyber Security: A Systematic Literature Review (2025). https://arxiv.org/abs/2405.04760

  50. [50]

    In: Proceedings of the Fourth ACM International Conference on AI in Finance

    Li, Y., Wang, S., Ding, H., Chen, H.: Large language models in finance: 22 A survey. In: Proceedings of the Fourth ACM International Conference on AI in Finance. ICAIF ’23, pp. 374–382. Association for Computing Machin- ery, New York, NY, USA (2023). https://doi.org/10.1145/3604237.3626869 . https://doi.org/10.1145/3604237.3626869

  51. [51]

    https://arxiv

    Yang, H., Wang, Y., Xu, X., Zhang, H., Bian, Y.: Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer (2024). https://arxiv. org/abs/2405.16856

  52. [52]

    Harvard Data Science Review7(1) (2025)

    Pawitan, Y., Holmes, C.: Confidence in the Reasoning of Large Language Models. Harvard Data Science Review7(1) (2025). https://hdsr.mitpress.mit.edu/pub/jaqt0vpb

  53. [53]

    https://arxiv.org/abs/2207.05221

    Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S....

  54. [54]

    Sub-Saharan Africa,

    Organization for Women in Science for the Developing World: Countries in the Global South. Accessed: 2024-11-23 (2020). https://owsd.net/sites/default/files/ OWSD%20138%20Countries%20-%20Global%20South.pdf Appendix A Supplementary Text A.1 Prompting Strategies and Response Evaluation We evaluated a total of nine models categorized into two groups based on...

  55. [56]

    Distinguish between factual claims and opinions: Determine if the statement can be objectively verified or if it’s a subjective view

  56. [57]

    Assess verifiability and plausibility: Check if the claim can be proven using reliable sources and if it aligns with established knowledge

  57. [58]

    Check for vague language, missing context, or unsupported generalizations: Identify if the claim lacks specificity or crucial details that could change its meaning

  58. [59]

    item Consider source credibility and potential biases: Evaluate the reliability of the claim’s origin and any possible motives behind it

  59. [60]

    Verify quantifiable elements against official statistics or studies: Cross-reference any numbers or data with trustworthy, authoritative sources

  60. [61]

    Be aware of logical fallacies and misleading rhetoric: Recognize common reasoning errors or manipulative language techniques

  61. [62]

    Evaluate time frame and currency of information: Determine if the claim is about current events or if it might be outdated

  62. [63]

    Break down complex claims into checkable components: Separate multi-part statements into individual, verifiable elements

  63. [64]

    Consider potential translation issues for non-native claims: Be mindful of possible misinterpretations if the claim was translated from another language

  64. [65]

    There is no credible evidence to support the claim

    Distinguish between causation and correlation: Identify if the claim incorrectly implies a cause-effect relationship when only a connection is shown. Guidelines: - Only output ‘True’ or ‘False’ when absolutely certain - Use ‘Other’ for ambiguous claims or insufficient information - Provide detailed explanation for your assessment Output format: Respond in...

  65. [2020]

    This is True

    Is this True or False? This is True. In May 2020, during protests against police brutality in Minneapolis, Min- nesota, law enforcement officers were seen slashing the tires of cars parked near the demonstrations. This was confirmed by local news reports and videos shared on social media. True The response is mapped to ‘True’ because it unequivocally and ...

  66. [2707]

    https: //doi.org/10.1145/3626772.3657914 .https://doi.org/10.1145/3626772.3657914

    Association for Computing Machinery, New York, NY, USA (2024). https: //doi.org/10.1145/3626772.3657914 .https://doi.org/10.1145/3626772.3657914