REVIEW 3 cited by
Assessing Large Language Models on Climate Information
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication.
Forward citations
Cited by 3 Pith papers
-
GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
GeoGrid-Bench evaluates 11 foundation models on 3,200 expert-curated questions about gridded climate data across 16 variables, finding vision-language models strongest and code generation weakest.
-
A RAG-Based Multi-Agent LLM System for Natural Hazard Resilience and Adaptation
WildfireGPT, a multi-agent RAG system with user profiling, outperforms ChatClimate and Perplexity AI in location-specific wildfire data analysis and evidence-based recommendations across ten expert case studies.
-
Challenges in Guardrailing Large Language Models for Science
A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.
Discussion (0). Continue with ORCID to comment.