Pith. sign in

REVIEW 8 cited by

ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16701 v2 pith:ICF3QNOF submitted 2024-10-22 cs.LG

ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

classification cs.LG
keywords climateevaluationframeworkllmscomprehensivedatasetdevelopissue
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientific validity of model outputs. To address this issue, we develop ClimaGen (Climate QA Generator), an adaptive learning framework that generates question-answer pairs from graduate textbooks with climate scientists in the loop. As a result, we present ClimaQA-Gold, an expert-annotated benchmark dataset alongside ClimaQA-Silver, a large-scale, comprehensive synthetic QA dataset for climate science. Finally, we develop evaluation strategies and compare different LLMs on our benchmarks. Our results offer novel insights into various approaches used to enhance knowledge of climate LLMs. The source code is publicly available at https://github.com/Rose-STL-Lab/genie-climaqa

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning

    cs.LG 2026-05 unverdicted novelty 7.0

    ReCrit frames critic interaction as a correctness-transition problem and uses quadrant-based RL rewards to improve LLM performance on scientific reasoning benchmarks by rewarding corrections and robustness while penal...

  2. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    Survey of RLM adoption in 28 disciplines reveals maturity disparities via a new assessment framework, with focus on development, evaluation, and public resources.

  3. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    A survey of RLM use in 28 disciplines reveals uneven adoption and introduces a maturity assessment framework showing larger gaps when limited to public resources.

  4. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 4.0

    A survey of reasoning language model adoption across 28 ERC scientific disciplines finds large maturity gaps, especially when only public resources are counted.

  5. OpenCompass: A Universal Evaluation Platform for Large Language Models

    cs.CL 2026-05 conditional novelty 4.0

    OpenCompass is a modular, high-concurrency platform for unified LLM evaluation across knowledge, reasoning, code, and other domains with support for rule-based, LLM-as-judge, and cascaded evaluators.

  6. OpenCompass: A Universal Evaluation Platform for Large Language Models

    cs.CL 2026-05 unverdicted novelty 3.0

    OpenCompass is presented as a one-stop, scalable, high-concurrency LLM evaluation platform with modular architecture supporting multiple domains and evaluator types.

  7. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 3.0

    The paper delivers a unified review and roadmap of Earth science foundation models, structured by capability depth from perception to agentic reasoning and by application breadth across atmosphere, hydrosphere, lithos...

  8. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 2.0

    A review of Earth science foundation models covering capability evolution from perception to discovery, applications across atmosphere/hydrosphere/lithosphere/biosphere/anthroposphere/cryosphere, over 200 datasets, an...