Pith. sign in

REVIEW 5 major objections 4 minor 64 references

Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking

T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read LLM-generated summaries of evidence let crowdsourced fact-checkers match the accuracy of full-length webpages while completing significantly more assessments in less time.

desk verdict Useful A/B result on LLM summaries for crowd truthfulness, but the 'comparable accuracy' claim needs equivalence testing before you trust it. read the letter →

arxiv 2501.18265 v2 pith:JMVADAU6 submitted 2025-01-30 cs.IR cs.CLcs.HC

classification cs.IRcs.CLcs.HC
keywords crowdsourcedfact-checkingtruthfulnessassessmentLLMsummarizationevidencePolitiFactcrowdsourcingefficiencyKrippendorff'salphaA/Btesting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish that LLM-generated summaries can replace full-length evidence documents in crowdsourced truthfulness assessment without loss of accuracy, while improving speed, cost, and worker agreement. It runs an A/B test on 120 PolitiFact statements: 100 U.S. crowd workers judged each statement with either full webpages (Standard) or bullet-point summaries produced by Meta-Llama-3-8B-Instruct (Summary). Accuracy and error metrics are statistically indistinguishable between the two conditions, and workers in the Summary condition complete about 15% more assessments in the same time, lowering projected costs at scale. The Summary condition also shows higher internal agreement among workers (Krippendorff's alpha), with no drop in reported reliance on or usefulness of the evidence. A sympathetic reader cares because this offers a practical, cheaper route to scaling fact-checking efforts at a moment when platforms are cutting professional fact-checking resources.

What carries the argument

The carrying mechanism is the Summary modality: each evidence webpage is condensed by Meta-Llama-3-8B-Instruct into a concise bullet-point list, guided by a prompt (Figure 1) that instructs the model to describe the document, focus on query-related points, stay accurate, and return JSON. The argument proceeds by comparing this modality against the full-length Standard condition on the same 120 PolitiFact statements, using accuracy, MSE, MAE, task time, Krippendorff's alpha, and self-reported need and usefulness as outcome measures. Summary fidelity is assumed rather than measured, so the comparison tests the summary as a presentation format, not the quality of the summarization itself.

What would settle it

Compare worker accuracy on a set of claims for which the LLM summary demonstrably omits a decisive detail, such as a quoted 'not' or a conditional caveat; if workers given those summaries systematically misjudge those claims compared to workers given full pages, then the comparable-effectiveness result depends on this particular summary quality rather than holding for summarized evidence generally.

Watch

Extended reading notes

Core claim

The central claim is that presenting crowd workers with LLM-generated summaries of evidence, rather than full-length webpages, yields comparable truthfulness assessments while significantly improving efficiency. In their experiment, individual accuracy is 0.31 in both modalities, and aggregated mean accuracy is 0.33 vs 0.31 for Standard vs Summary; bootstrap confidence intervals for the differences in MAE and MSE include zero, so no statistically significant gap appears. Workers in the Summary modality show higher agreement (Krippendorff's alpha) at p<0.05, and they complete roughly 15% more assessments per unit time. The paper reads these results as evidence that summarization preserves the decision-relevant content of the evidence while reducing cognitive load, making it a viable substitute for full-length evidence in large-scale truthfulness annotation.

Load-bearing premise

The load-bearing premise is that the LLM summaries preserve the factual content needed to judge truthfulness; the paper asserts this design goal in the prompt but never validates it against source documents or another summarizer.

Editorial extensions

If this is right

  • Crowdsourced fact-checking pipelines can swap full evidence pages for LLM summaries without losing accuracy, shrinking per-task time and cost.
  • Higher internal agreement in the Summary condition implies that condensed evidence reduces interpretational variance, which should make aggregated verdicts more stable.
  • Comparable over- and under-estimation error patterns across modalities indicate summaries do not systematically bias truthfulness judgments in one direction.
  • Efficiency gains compound with scale: the paper estimates 6,000 assessments cost about £1,716 with full pages but £1,486 with summaries, a saving that grows with dataset size.
  • Worker reliance on and perceived usefulness of evidence stay high with summaries, so the format does not cause workers to skip the evidence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because summary fidelity is never validated, the comparable-accuracy result may be specific to this Llama-3 prompt and model on PolitiFact-style webpages; a summarizer that drops caveats or flips negations could break the equivalence. A testable extension: run the same A/B with an adversarial summarizer that deletes key qualifiers and check whether accuracy diverges.
  • The ~15% throughput gain likely understates real-world gains, since the Summary condition also had lower initial abandonment (38% vs 49%); combining summaries with per-judgment payment could further cut effective cost per completed label.
  • Higher internal agreement in the Summary condition could reflect reduced information rather than better judgment: if summaries make workers converge on the same (possibly wrong) reading, agreement rises without accuracy improving. A follow-up could compare agreement on summaries known to omit a decisive fact.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper reports an A/B crowdsourcing study comparing two evidence-presentation modalities for truthfulness assessment: a Standard condition showing full-length webpages and a Summary condition showing LLM-generated (Meta-Llama-3-8B-Instruct) summaries of those webpages. Using 120 PolitiFact statements and 100 Prolific workers per condition, the authors measure agreement with expert labels (accuracy, MAE, MSE), internal agreement, task completion time, and worker-reported reliance on and usefulness of evidence. The paper claims that the Summary modality achieves accuracy and error metrics comparable to the Standard modality while significantly improving efficiency and lowering cost, and that it also increases internal agreement among workers.

Significance. If the central claim were established, the paper would make a useful practical contribution: LLM-generated summaries could reduce the time and cost of large-scale crowdsourced fact-checking without degrading judgment quality. The study has genuine strengths: it uses a real crowdsourcing platform, a controlled A/B design, external expert labels from PolitiFact, and it makes the collected data publicly available. The efficiency result is supported by a significant Mann-Whitney test, and the cost simulation, once corrected, gives a plausible order-of-magnitude estimate. However, the headline effectiveness claim rests on interpreting non-significant differences as equivalence, with wide confidence intervals and no pre-specified equivalence margin. The paper also does not validate whether the summaries preserve the factual content needed for truthfulness judgments. These issues are load-bearing for the paper's main conclusion, so the contribution is currently conditional on additional statistical and fidelity analysis.

major comments (5)
  1. [4.1] In Section 4.1 the authors state that because bootstrapped 95% confidence intervals for the MAE difference (−0.167 to 0.112) and MSE difference (−0.767 to 0.340) include zero, there is no statistically significant difference, and then conclude that the modalities are comparable. This is the central claim of RQ1, but non-significance is not evidence of equivalence; on the six-level truthfulness scale, an MAE difference of 0.112 is a meaningful degradation. The paper needs a pre-specified equivalence margin, a two one-sided test (TOST), or, at minimum, an effect size with a power analysis to justify 'comparable.' Without this, the headline effectiveness conclusion is underdetermined.
  2. [5.2] Section 5.2 repeats the same logical gap: Table 4 and the simulation in Figure 6 are used to conclude comparable performance under equal time constraints because 'no statistical significance is detected.' The same equivalence-margin requirement applies. Additionally, Section 5.2's judgment multiplier is +15%, yet Section 7 states that workers 'complete nearly double the number of judgments within the same time frame.' These statements are inconsistent, and the conclusions section should be reconciled with Table 4.
  3. [3.2] The summary-fidelity premise is asserted but never validated. The prompt in Figure 1 instructs the LLM to produce query-focused summaries, and Section 3.2 states the design goal of preserving factuality and stance, but the paper provides no evaluation of whether the summaries omit key caveats or introduce distortions, no comparison with another summarization model or prompt, and no error analysis of summaries. Because the general claim is that LLM summaries can replace full evidence, the absence of any fidelity check leaves open that the observed comparability, if real, is an artifact of this particular model and prompt. At minimum, a sample-level manual fidelity audit or a summary-evidence entailment check is needed.
  4. [5.1] The cost simulation in Section 5.1 contains a unit inconsistency. The authors state that 600 assessments require 23.67 hours at a U.S. minimum wage of $7.25/hour (≈ £6.08), and report costs of £171.60 (Standard) and £148.63 (Summary). These figures equal 23.67 × 7.25 and 20.50 × 7.25, i.e., dollar amounts, not pounds. The correct pound figures at £6.08/hour would be approximately £143.9 and £124.6. This error affects the quantitative cost-savings claim and must be corrected.
  5. [4.3] Table 3 contains internally inconsistent count/percentage pairs. In the Standard False row, the counts 39, 44, 35 sum to 118, but the percentages sum to 100% and correspond to different denominators; in the Half True rows the counts sum to 100 but the percentages (48.5%, 30.1%, 20.6%) do not match those counts. The table therefore cannot be used to support the claim that error-type distributions are similar across modalities, and it needs to be recomputed or clearly explained.
minor comments (4)
  1. [4.2] The Mann-Whitney U test comparing Krippendorff's α values is performed on a very small set of aggregate values (six ground-truth levels plus overall); the paper should state the sample size and justify treating these values as independent observations, or use a worker-level bootstrap for agreement.
  2. [6.3] The notation 'σ2 = 0.67' and 'σ2 = 0.99' is inconsistent with the text's use of standard deviation elsewhere; please clarify whether these are variances or standard deviations.
  3. [6.1] The sentence 'In theStandard modality, there is a notable pattern where evidence is frequently considered at least moderately useful, which aligns with the reduced amount of information presented to the workers' seems to describe the Summary condition rather than the Standard condition; please rephrase.
  4. [3.3] The minimum time requirement of 3 seconds per statement is extremely low and could admit speeders; please report how many workers were excluded by this check and by the gold-standard questions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the summary-versus-standard comparison is an empirical A/B study benchmarked against external PolitiFact expert labels, and the central claim is not assumed by the inputs.

full rationale

The paper is an empirical A/B comparison in which the central claim—that Summary evidence yields comparable accuracy and error metrics to Standard evidence while improving efficiency—is evaluated against external PolitiFact expert labels and observed worker judgments. No parameter is fitted to the outcome and then re-reported as a prediction; the comparison metrics (accuracy, MAE, MSE, time) are computed directly from crowd responses and ground truth. The reuse of the La Barbera et al. dataset and task design is inheritance of instrumentation, not circular reasoning: the prior work supplies the task layout and six-level scale, but the target result (summary versus full evidence) is not assumed by those inputs. The summarization prompt in Figure 1 is an LLM configuration choice, not a derivation step, and the paper explicitly acknowledges in the Conclusions that automatically generated summaries may omit critical nuances. Concerns about interpreting non-significant bootstrapped confidence intervals as equivalence (Sections 4.1 and 5.2) are statistical-evidence concerns, not circularity: 'no statistically significant difference' is not an input that produces the conclusion by construction. No self-citation chain forces the main finding, and no known result is merely renamed or repackaged. Hence no significant circularity is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim depends on several unverified domain assumptions: expert PolitiFact labels are ground truth, the Llama summaries preserve factually relevant content, and non-significant differences are treated as comparability. In addition, the 3-second minimum-time filter and the different per-arm compensation are hand-set design parameters whose effects are not reported. No new physical or theoretical entities are introduced.

free parameters (3)
  • Minimum time threshold = 3 seconds per statement
    Chosen by hand as a quality filter; the number of workers excluded is not reported, so its effect on the compared pools is unknown.
  • Compensation rates = £2.00 Standard, £1.80 Summary
    Different pay per arm is a design parameter that can affect speed, effort, and dropout; it is not accounted for in the efficiency comparison beyond the headline wage simulation.
  • Summarization prompt wording = Prompt in Figure 1
    The prompt was selected after several rounds of trial and error, and no ablation is given, so the results are tied to this specific hand-tuned prompt.
assumptions (5)
  • domain assumption PolitiFact expert labels are treated as ground truth for accuracy and error metrics.
    Accuracy and error are computed against expert labels; if those labels are noisy or biased, the reported effectiveness numbers inherit that noise. Invoked throughout Section 4.
  • domain assumption LLM-generated summaries preserve the factual content needed for truthfulness assessment.
    The prompt in Figure 1 asks for accurate summaries, but the paper does not validate summary fidelity against source documents. This assumption is load-bearing for the effectiveness comparison.
  • ad hoc to paper Absence of a statistically significant difference is interpreted as comparable effectiveness.
    The paper concludes "comparable" from bootstrap confidence intervals that include zero, without a pre-specified equivalence margin or power analysis. This interpretive step is introduced by the authors, not established by the data.
  • domain assumption Workers passing gold-standard and 3-second checks are attentive and comparable across the two arms.
    The quality filters in Section 3.3 are standard, but exclusion counts are not reported, and the two arms are paid differently, so differential attrition or motivation cannot be ruled out.
  • domain assumption Bing-retrieved webpages are relevant evidence for each claim.
    The dataset includes at least 10 retrieved webpages per claim, but retrieval quality is not evaluated. Poor retrieval would affect both modalities, possibly masking interactions with summarization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking." pith.science (2026). https://pith.science/paper/JMVADAU6

@misc{pith2026250118265,
  author       = {Pith},
  title        = {Pith review of: Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JMVADAU6}},
  note         = {Machine review of arXiv:2501.18265}
}
read the original abstract

Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with a large language model. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions. Our analysis explores both the quality of assessments and the behavioral patterns of participants. The results reveal that relying on summarized evidence offers comparable accuracy and error metrics to the Standard modality while significantly improving efficiency. Workers in the Summary setting complete a significantly higher number of assessments, reducing task duration and costs. Additionally, the Summary modality maximizes internal agreement and maintains consistent reliance on and perceived usefulness of evidence, demonstrating its potential to streamline large-scale truthfulness evaluations.

Figures

Figures reproduced from arXiv: 2501.18265 by the authors.

Figure 1
Figure 1. Evidence summarization prompt. The process begins with a carefully crafted prompt provided to the model. This prompt is specifically designed to guide the model in extracting and summarizing the essential factual elements from the content of a webpage, ensuring that the generated summary maintains consistency with the original information in terms of factuality and stance. After several rounds of trial and error, th… view at source ↗
Figure 2
Figure 2. Agreement between workers and experts on individual and aggregated assessments. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Krippendorff’s α score [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Accuracy across sampling sizes. 4.3 Error Analysis To better understand the patterns of errors in truthfulness assessments, we conducted a failure analysis by categorizing worker assessments into three types: correct, overestimation, and underestimation. The results, s…
Figure 5
Figure 5. Figure 5: Time elapsed on statements for the two modalities. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Distributions of Accuracy, Mean Squared Error (MSE), and Mean Absolute Error (MAE) [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Need and perceived usefulness of the evidence in the task. [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 31 canonical work pages

  1. [1]

    Alim Al Ayub Ahmed, Ayman Aljabouh, Praveen Kumar Donepudi, and Myung Suh Choi. 2021. Detecting Fake News Using Machine Learning: A Systematic Literature Review.Psychology and Education Journal58, 1 (2021), 10 pages. https://doi.org/10.17762/pae.v58i1.1046

  2. [2]

    Jennifer Allen, Cameron Martel, and David G Rand. 2022. Birds of a Feather Don’t Fact-Check Each Other: Partisanship and the Evaluation of News in Twitter’s Birdwatch Crowdsourced Fact- Checking Program. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New ...

  3. [3]

    Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. Generating Fact Checking Explanations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 7352–7364.https://do...

  4. [4]

    Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2022. Fact Checking with Insufficient Evidence.Transactions of the Association for Computational Linguistics 10 (2022), 746–763. https://doi.org/10.1162/tacl_a_00486

  5. [5]

    Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, Eduard Hovy, Heng Ji, Filippo Menczer, Ruben Miguez, Preslav Nakov, Dietram Scheufele, Shivam Sharma, and Giovanni Zagni. 2024. Factuality challenges in the era of large language models and...

  6. [6]

    Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference o...

  7. [7]

    Nadav Borenstein, Greta Warren, Desmond Elliott, and Isabelle Augenstein. 2025. Can Community Notes Replace Professional Fact-Checkers? arXiv. https://doi.org/10.48550/arXiv.2502. 14132 arXiv:2502.14132 [cs.CL]

  8. [8]

    Erik Brand, Kevin Roitero, Michael Soprano, Afshin Rahimi, and Gianluca Demartini. 2022. A Neural Model to Jointly Predict and Explain Truthfulness of Statements.Journal of Data and Information Quality 15, 1, Article 4 (12 2022), 19 pages.https://doi.org/10.1145/3546917

Show all 64 references
  1. [9]

    Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, and Gianluca Demartini

  2. [10]

    Botambu Collins, Dinh Tuyen Hoang, Ngoc Thanh Nguyen, and Dosam Hwang. 2021. Trends in Combating Fake News On Social Media – A Survey.Journal of Information and Telecommunication 5, 2 (2021), 247–266. https://doi.org/10.1080/24751839.2020.1847379

  3. [11]

    Alphaeus Dmonte, Roland Oruche, Marcos Zampieri, Prasad Calyam, and Isabelle Augenstein

  4. [12]

    Xishuang Dong, Shouvon Sarker, and Lijun Qian. 2022. Integrating Human-in-the-loop into Swarm Learning for Decentralized Fake News Detection. InInternational Conference on Intelligent Data Science Technologies and Applications (IDSTA). IEEE, San Antonio, TX, USA, 46–53. https:...

  5. [13]

    Tim Draws, David La Barbera, Michael Soprano, Kevin Roitero, Davide Ceolin, Alessandro Checco, and Stefano Mizzaro. 2022. The Effects of Crowd Worker Biases in Fact-Checking Tasks. In 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association fo...

  6. [14]

    Tim Draws, Alisa Rieger, Oana Inel, Ujwal Gadiraju, and Nava Tintarev. 2021. A Checklist to Combat Cognitive Biases in Crowdsourcing. InProceedings of the Ninth AAAI Conference on Human Computation and Crowdsourcing. AAAI Press, Palo Alto, California, USA, 48–59. https://doi.o...

  7. [15]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models. arXiv. https://doi.org/10.48550/arXiv.2407.21783 arXiv:2407.21783

  8. [16]

    Tibshirani

    Bradley Efron and Robert J. Tibshirani. 1993.An Introduction to the Bootstrap. Monographs on Statistics and Applied Probability, Vol. 57. Chapman & Hall/CRC, New York. https: //doi.org/10.1201/9780429246593

  9. [17]

    Carsten Eickhoff. 2018. Cognitive Biases in Crowdsourcing. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining(Marina Del Rey, CA, USA) (WSDM ’18). Association for Computing Machinery, New York, NY, USA, 162–170.https: //doi.org/10.1145/31...

  10. [18]

    Nicola Ferro, Yubin Kim, and Mark Sanderson. 2019. Using Collection Shards to Study Retrieval Performance Effect Sizes.ACM Transactions on Information Systems37, 3, Article 30 (March 2019), 40 pages. https://doi.org/10.1145/3310364

  11. [19]

    Nicola Ferro and Gianmaria Silvello. 2016. A General Linear Mixed Models Approach to Study System Component Effects. InProceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval(Pisa, Italy)(SIGIR ’16). Association for Compu...

  12. [20]

    Nicola Ferro and Gianmaria Silvello. 2018. Toward an anatomy of IR system component perfor- mances. Journal of the Association for Information Science and Technology69, 2 (2018), 187–200. https://doi.org/10.1002/asi.23910

  13. [21]

    Freedman

    David A. Freedman. 2009.Statistical Models: Theory and Practice(2 ed.). Cambridge University Press, New York, NY, USA. https://doi.org/10.1017/CBO9780511815867

  14. [22]

    Meric Altug Gemalmaz and Ming Yin. 2021. Accounting for Confirmation Bias in Crowdsourced Label Aggregation. In Proceedings of the Thirtieth International Joint Conference on Artifi- cial Intelligence (IJCAI). International Joint Conferences on Artificial Intelligence Organiza...

  15. [23]

    Lei Han, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, and Gianluca Demartini. 2019. All Those Wasted Hours: On Task Abandonment in Crowdsourcing. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (Me...

  16. [24]

    Linmei Hu, Siqi Wei, Ziwang Zhao, and Bin Wu. 2022. Deep Learning For Fake News Detection: A Comprehensive Survey.AI Open3 (2022), 133–155. https://doi.org/10.1016/j.aiopen. 2022.09.001

  17. [25]

    Christoph Hube, Besnik Fetahu, and Ujwal Gadiraju. 2019. Understanding and Mitigating Worker Biases in the Crowdsourced Collection of Subjective Judgments. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk)(CHI ’19). Associatio...

  18. [26]

    Klaus Krippendorff. 2011. Computing Krippendorff’s Alpha-Reliability.https://repository. upenn.edu/asc_papers/43. Technical Report, University of Pennsylvania

  19. [27]

    David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, and Stefano Mizzaro. 2024. Crowdsourced Fact-checking: Does It Actually Work? Information Processing & Management 61, 5 (2024), 103792. https: //doi.org/10.1016/...

  20. [28]

    David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, and Damiano Spina

  21. [29]

    David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonatha...

  22. [30]

    Huiru Li, Liangxiao Jiang, and Siqing Xue. 2023. Neighborhood Weighted Voting-Based Noise Correction for Crowdsourcing. ACM Transactions on Knowledge Discovery from Data17, 7, Article 96 (4 2023), 18 pages. https://doi.org/10.1145/3586998

  23. [31]

    Houjiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou, Daisy Pinaroc, Matthew Lease, and Min Kyung Lee. 2024. Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI.Proceedings of the ACM on Human-Computer Interaction8, CSCW2, Article 423 (...

  24. [32]

    Houjiang Liu, Jacek Gwizdka, and Matthew Lease. 2024. Exploring Multidimensional Check- worthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers. arXiv.https: //doi.org/10.48550/arXiv.2412.08185

  25. [33]

    Mann and Donald R

    Henry B. Mann and Donald R. Whitney. 1947. On a Test of Whether One of Two Random Variables is Stochastically Larger than the Other.Annals of Mathematical Statistics18, 1 (1947), 50–60. https://doi.org/10.1214/aoms/1177730491

  26. [34]

    Syed Ishfaq Manzoor, Jimmy Singla, and Nikita. 2019. Fake News Detection Using Machine Learning approaches: A systematic Review. InProceedings of the 3rd International Conference on Trends in Electronics and Informatics (ICOEI). IEEE, Tirunelveli, India, 230–234. https: //doi....

  27. [35]

    Cameron Martel, Jennifer Allen, Gordon Pennycook, and David G. Rand. 2024. Crowds Can Effectively Identify Misinformation at Scale.Perspectives on Psychological Science19, 2 (2024), 477–488. https://doi.org/10.1177/17456916231190388

  28. [36]

    Preslav Nakov, Alberto Barrón-Cedeño, Giovanni da San Martino, Firoj Alam, Julia Maria Struß, Thomas Mandl, Rubén Míguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Seli...

  29. [37]

    Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barrón-Cedeño, Rubén Míguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Struß, and Thomas Mandl. 2021. The CLEF-2021 CheckThat! Lab o...

  30. [38]

    Thanh Tam Nguyen, Matthias Weidlich, Hongzhi Yin, Bolong Zheng, Quang Huy Nguyen, and Quoc Viet Hung Nguyen. 2020. FactCatch: Incremental Pay-as-You-Go Fact Checking with Minimal User Effort. InProceedings of the 43rd International ACM SIGIR Conference on Research and Developm...

  31. [39]

    Olejnik and James Algina

    Stephen F. Olejnik and James Algina. 2003. Generalized Eta and Omega Squared Statistics: Measures of Effect Size for Some Common Research Designs.Psychological Methods8, 4 (2003), 434–447. https://doi.org/10.1037/1082-989X.8.4.434

  32. [40]

    Gordon Pennycook and David G. Rand. 2019. Fighting misinformation on social media using crowdsourced judgments of news source quality.Proceedings of the National Academy of Sciences 116, 7 (2019), 2521–2526. https://doi.org/10.1073/pnas.1806781116

  33. [41]

    Gordon Pennycook and David G. Rand. 2021. The Psychology of Fake News.Trends in Cognitive Sciences 25, 5 (2021), 388–402. https://doi.org/10.1016/j.tics.2021.02.007

  34. [42]

    Yunke Qu, Kevin Roitero, David La Barbera, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2022. Combining Human and Machine Confidence in Truthfulness Assessment.Journal of Data and Information Quality15, 1, Article 5 (dec 2022), 17 pages.https://doi.org/10. 1145/3546916

  35. [43]

    Benjamin Riedel, Isabelle Augenstein, Georgios P Spithourakis, and Sebastian Riedel. 2017. A simple but tough-to-beat baseline for the fake news challenge stance detection task. arXiv. https://doi.org/10.48550/arXiv.1707.03264

  36. [44]

    LeveragingBehavioral Heterogeneity Across Markets for Cross-Market Training of Recommender Systems

    KevinRoitero, BenCarterette, RishabhMehrotra, andMouniaLalmas.2020. LeveragingBehavioral Heterogeneity Across Markets for Cross-Market Training of Recommender Systems. InCompanion Proceedings of the Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machin...

  37. [45]

    Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2020. Can The Crowd Identify Misinformation Objectively? The Effects of Judgment Scale and Assessor’s Background. InProceedings of the 43rd International ACM SIGIR Conference ...

  38. [46]

    Mohammed Saeed, Nicolas Traub, Maelle Nicolas, Gianluca Demartini, and Paolo Papotti. 2022. Crowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management (Atlanta...

  39. [47]

    Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Automated Verification of Textual Claims (AVeriTeC) Shared Tas...

  40. [48]

    Vinay Setty. 2024. Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models. InProceedings of the 47th International ACM SIGIR Conference on Research 17 and Development in Information Retrieval(Washington DC, USA)(SIGIR ’24). Association for...

  41. [49]

    Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Gian- luca Demartini, and Stefano Mizzaro. 2024. Cognitive Biases in Fact-Checking and Their Countermeasures: A Review. Information Processing & Management 61, 3 (2024), 103672. https://doi.org/10....

  42. [50]

    Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2021. The Many Dimensions of Truthfulness: Crowdsourcing Misinformation Assessments on a Multidimensional Scale.Information Processing & Management 58, 6 (2...

  43. [51]

    All Abdullah Tanvir, Ehesas Mia Mahir, Saima Akhter, and Mohammad Rezwanul Huq. 2019. Detecting Fake News using Machine Learning and Deep Learning Algorithms. InProceedings of the 7th International Conference on Smart Computing & Communications (ICSCC). IEEE, Piscataway, NJ, U...

  44. [52]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a Large-scale Dataset for Fact Extraction and VERification. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...

  45. [53]

    James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal

  46. [54]

    Nikhita Vedula and Srinivasan Parthasarathy. 2021. FACE-KEG: Fact Checking Explained using KnowledgE Graphs. InProceedings of the 14th ACM International Conference on Web Search and Data Mining (Virtual Event, Israel)(WSDM ’21). Association for Computing Machinery, New York, N...

  47. [56]

    Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Ruba- shevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, and Preslav Nakov. 2024. Factcheck-Bench: Fine- Grained Evalua...

  48. [57]

    Greta Warren, Irina Shklovski, and Isabelle Augenstein. 2025. Show Me the Work: Fact-Checkers’ RequirementsforExplainableAutomatedFact-Checking.In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM, 1–21. https://doi.org/10.1145/ 370659...

  49. [58]

    Jing Yang, Didier Vega-Oliveros, Tais Seibt, and Anderson Rocha. 2021. Scalable Fact-checking with Human-in-the-Loop. In2021 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, Montpellier, France, 1–6.https://doi.org/10.1109/WIFS53200.2021.9648388 18

  50. [59]

    Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang. 2023. End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inform...

  51. [60]

    Shane Culpepper, Oren Kurland, and Stefano Mizzaro

    Fabio Zampieri, Kevin Roitero, J. Shane Culpepper, Oren Kurland, and Stefano Mizzaro. 2019. On Topic Difficulty in IR Evaluation: The Effect of Systems, Corpora, and System Components. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in...

  52. [61]

    Andy Zhao and Mor Naaman. 2023. Variety, Velocity, Veracity, and Viability: Evaluating the Contributions of Crowdsourced and Professional Fact-checking.https://doi.org/10.31235/ osf.io/yfxd3 19

  53. [2017]

    Proceedings of the AAAI Conference on Human Computation and Crowdsourcing5, 1 (Sep

    Let’s Agree to Disagree: Fixing Agreement Measures for Crowdsourcing. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing5, 1 (Sep. 2017), 11–20. https://doi.org/10.1609/hcomp.v5i1.13306

  54. [2019]

    The FEVER2.0 Shared Task. In Proceedings of the Second Workshop on Fact Extrac- tion and VERification (FEVER), James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal (Eds.). Association for Computational Linguistics, Hong Kong, China, 1–6....

  55. [2020]

    InAdvances in Information Retrieval, Joemon M

    Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias. InAdvances in Information Retrieval, Joemon M. Jose, Emine Yilmaz, João Magalhães, Pablo Castells, Nicola Ferro, Mário J. Silva, and Flávio Martins (Eds.). Springer International Publishing, Cham, 207–...

  56. [2025]

    Claim Verification in the Age of Large Language Models: A Survey. arXiv. https: //doi.org/10.48550/arXiv.2408.14317 arXiv:2408.14317 [cs.CL]

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.