Climate Finance Bench releases 330 expert-validated QA pairs on 33 climate reports and shows that retrieval quality, not model capacity, is the main accuracy bottleneck.
ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
To handle the vast amounts of qualitative data produced in corporate climate communication, stakeholders increasingly rely on Retrieval Augmented Generation (RAG) systems. However, a significant gap remains in evaluating domain-specific information retrieval - the basis for answer generation. To address this challenge, this work simulates the typical tasks of a sustainability analyst by examining 30 sustainability reports with 16 detailed climate-related questions. As a result, we obtain a dataset with over 8.5K unique question-source-answer pairs labeled by different levels of relevance. Furthermore, we develop a use case with the dataset to investigate the integration of expert knowledge into information retrieval with embeddings. Although we show that incorporating expert knowledge works, we also outline the critical limitations of embeddings in knowledge-intensive downstream domains like climate change communication.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
baseline 1representative citing papers
citing papers explorer
-
Climate Finance Bench
Climate Finance Bench releases 330 expert-validated QA pairs on 33 climate reports and shows that retrieval quality, not model capacity, is the main accuracy bottleneck.