REVIEW 3 cited by
Evaluating open-source Large Language Models for automated fact-checking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The increasing prevalence of online misinformation has heightened the demand for automated fact-checking solutions. Large Language Models (LLMs) have emerged as potential tools for assisting in this task, but their effectiveness remains uncertain. This study evaluates the fact-checking capabilities of various open-source LLMs, focusing on their ability to assess claims with different levels of contextual information. We conduct three key experiments: (1) evaluating whether LLMs can identify the semantic relationship between a claim and a fact-checking article, (2) assessing models' accuracy in verifying claims when given a related fact-checking article, and (3) testing LLMs' fact-checking abilities when leveraging data from external knowledge sources such as Google and Wikipedia. Our results indicate that LLMs perform well in identifying claim-article connections and verifying fact-checked stories but struggle with confirming factual news, where they are outperformed by traditional fine-tuned models such as RoBERTa. Additionally, the introduction of external knowledge does not significantly enhance LLMs' performance, calling for more tailored approaches. Our findings highlight both the potential and limitations of LLMs in automated fact-checking, emphasizing the need for further refinements before they can reliably replace human fact-checkers.
Forward citations
Cited by 3 Pith papers
-
Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Post-cut-off MAFC claims remain 17–29% potentially contaminated; contamination inflates Macro-F1 by up to 11.34 points and can change model rankings.
-
ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation
ArgRAG builds a weighted bipolar argumentation graph from retrieved documents, computes evidence strengths with quadratic energy semantics, and classifies claims by the final strength of the claim node.
-
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
Chinese search engines, LLMs, and AI Overviews are comparably accurate when they commit to an answer but differ sharply in answer rate, and all are systematically better on yes-labeled than no-labeled health-dominated...
Discussion (0). Sign in to comment.