Pith. sign in

REVIEW 16 cited by

CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.00614 v2 pith:KPLRRMHG submitted 2020-12-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords claimsclimatedatasetclimate-fevercommunityfeverlanguagereal-world
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce CLIMATE-FEVER, a new publicly available dataset for verification of climate change-related claims. By providing a dataset for the research community, we aim to facilitate and encourage work on improving algorithms for retrieving evidential support for climate-specific claims, addressing the underlying language understanding challenges, and ultimately help alleviate the impact of misinformation on climate change. We adapt the methodology of FEVER [1], the largest dataset of artificially designed claims, to real-life claims collected from the Internet. While during this process, we could rely on the expertise of renowned climate scientists, it turned out to be no easy task. We discuss the surprising, subtle complexity of modeling real-world climate-related claims within the \textsc{fever} framework, which we believe provides a valuable challenge for general natural language understanding. We hope that our work will mark the beginning of a new exciting long-term joint effort by the climate science and AI community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mutual Linearity in and out of Stationarity for Markov Jump Processes: A Trajectory-Based Approach

    cond-mat.stat-mech 2026-04 unverdicted novelty 7.0 of 10

    A trajectory-level derivation shows mutual linearity holds for non-stationary Markov jump processes and generalizes to other systems.

  2. Tailored untruths: How personalisation challenges LLM safeguards

    cs.CL 2025-10 conditional novelty 7.0 of 10

    A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.

  3. GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

    cs.CL 2024-12 unverdicted novelty 7.0 of 10

    GME achieves state-of-the-art results in universal multimodal retrieval by training on a balanced synthetic multimodal dataset.

  4. Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

    physics.soc-ph 2026-07 unverdicted novelty 6.0 of 10

    LLMs show three distinct non-sycophantic responses to science skepticism, with robustness in some cases being accidental because the model does not represent the skepticism signal, as determined by linear probes on th...

  5. Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models

    cs.IR 2026-04 conditional novelty 6.0 of 10

    RouteHead trains a lightweight router to dynamically select optimal LLM attention heads per query for improved attention-based document re-ranking.

  6. Data, Not Model: Explaining Bias toward LLM Texts in Neural Retrievers

    cs.IR 2026-04 unverdicted novelty 6.0 of 10

    Bias toward LLM texts in neural retrievers arises from artifact imbalances between positive and negative documents in training data that are absorbed during contrastive learning.

  7. Score-Only Distillation for Compact Dense Retrieval

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Score-only distillation with a row-centered all-pairs PairMSE objective lets 0.6B bi-encoders recover up to 50% of the base-to-teacher retrieval gap under matched protocols.

  8. Uncertainty-Aware Web-Conditioned Scientific Fact-Checking

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    An uncertainty-gated fact-checking system decomposes claims atomically, verifies them against context, and selectively searches the web only for uncertain facts, outperforming benchmarks while abstaining on conflicts.

  9. Mutual Linearity in and out of Stationarity for Markov Jump Processes: A Trajectory-Based Approach

    cond-mat.stat-mech 2026-04 unverdicted novelty 5.0 of 10

    Trajectory-level linear response yields mutual linearity of observables under single-edge rate perturbation for Markov jump processes, including non-stationary state and counting observables.

  10. Text Embeddings by Weakly-Supervised Contrastive Pre-training

    cs.CL 2022-12 unverdicted novelty 5.0 of 10

    E5 text embeddings trained with weakly-supervised contrastive pre-training on CCPairs outperform BM25 on BEIR zero-shot and achieve top results on MTEB, beating much larger models.

  11. Evidence-Ledger Adjudication for Claim-Evidence Traceability

    cs.AI 2026-07 conditional novelty 4.0 of 10

    An evidence-ledger workflow labels claim-evidence pairs as supported/contradicted/missing/mixed and routes unsupported claims back to authors, reporting 0.676 accuracy over TF-IDF's 0.383 on a 2,335-row benchmark.

  12. Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models

    cs.AI 2026-04 unverdicted novelty 4.0 of 10

    DAVinCI combines claim attribution to model internals and external sources with entailment-based verification to improve LLM factual reliability by 5-20% on fact-checking datasets.

  13. HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers

    cs.IR 2025-09 conditional novelty 4.0 of 10

    By first fusing multiple retrievers within labeled and unlabeled sources with RRF, then merging z-score normalized lists, HF-RAG improves fact-verification F1 in-domain and out-of-domain.

  14. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 3.0 of 10

    The paper delivers a unified review and roadmap of Earth science foundation models, structured by capability depth from perception to agentic reasoning and by application breadth across atmosphere, hydrosphere, lithos...

  15. Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrieval

    cs.IR 2026-04 conditional novelty 3.0 of 10

    Reproducibility study confirms Hypencoder's non-linear query-specific scoring improves retrieval over bi-encoders on standard benchmarks but standard methods remain faster and hard-task results are mixed due to implem...

  16. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 2.0 of 10

    A review of Earth science foundation models covering capability evolution from perception to discovery, applications across atmosphere/hydrosphere/lithosphere/biosphere/anthroposphere/cryosphere, over 200 datasets, an...

Pith tools