REVIEW 5 cited by
Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While large pre-trained language models (LLMs) have shown their impressive capabilities in various NLP tasks, they are still under-explored in the misinformation domain. In this paper, we examine LLMs with in-context learning (ICL) for news claim verification, and find that only with 4-shot demonstration examples, the performance of several prompting methods can be comparable with previous supervised models. To further boost performance, we introduce a Hierarchical Step-by-Step (HiSS) prompting method which directs LLMs to separate a claim into several subclaims and then verify each of them via multiple questions-answering steps progressively. Experiment results on two public misinformation datasets show that HiSS prompting outperforms state-of-the-art fully-supervised approach and strong few-shot ICL-enabled baselines.
Forward citations
Cited by 5 Pith papers
-
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
On deception detection benchmarks, fine-tuned transformers beat LLMs on data-rich datasets, few-shot GPT-4o wins the small legal corpus, and chain-of-thought prompting frequently reduces F1.
-
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
REFLEX improves explainable fact-checking by using verdict-anchored style control and self-disagreement signals to disentangle fact from style in LLM outputs, achieving SOTA results with minimal self-refined samples.
-
Recon, Answer, Verify: Agents in Search of Truth
Removing annotator cues from fact-checking evidence lowers LLM scores substantially, and a three-agent question-answering pipeline, RAV, outperforms several published fact-checking baselines.
-
Multimodal rumor detection enhanced by external evidence and forgery features
A combination of Fourier forgery features, BLIP captions, evidence attention, and gated fusion reports 94.9% macro accuracy on Weibo and 94.1% on Twitter for rumor detection on MR2, but the closest baseline is omitted...
-
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.
Discussion (0). Sign in to comment.