Pith. sign in

Paper Citation Record · LEDGER

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2505.17139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17139 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:21.111443Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:22:23.855247Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:22:26.516274Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12e6ab02-ab96-4435-9f59-007f7f320169 · outbound

This paper cites GPT-4 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.566305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.566305Z digest=sha256:a3a06a48692e5acbda42f117d6af67159cc44854e93db3b71198fd06b44794fc

Observation 6f7f9f8a-0dc7-4e7c-b5ab-106c5b6bd07a · outbound

This paper cites OceanGPT: A Large Language Model for Ocean Science Tasks.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OceanGPT: A Large Language Model for Ocean Science Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.576960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.576960Z digest=sha256:e9376b8711e374d29c7eb4669a423dcb2c43812b5a3554bb857d43dd2576e732

Observation 3b9a4a9b-84d1-49dc-8890-d6f090abcec9 · outbound

This paper cites Matplotlib and seaborn.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Matplotlib and seaborn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.734997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.586447Z digest=sha256:5246dd115c8b7ff820a5bdfe6c5d939eb649cc765c9cdf381cc9bf8343dcc482

Observation ee12df61-3140-427d-89bd-3baa7c2563f7 · outbound

This paper cites This reference does not exist: an exploration of llm citation accuracy and relevance.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs This reference does not exist: an exploration of llm citation accuracy and relevance

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.704132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.597692Z digest=sha256:c23a96aaba38b4ff226bde89f8ed8a82ed0c940ea9478fbc67e219342c707aa5

Observation 03694014-52e2-4d23-a59a-ca3afde7f50d · outbound

This paper cites SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.605939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.605939Z digest=sha256:801ac449655c821aee7f8349535f71b486bf3c6595d3cb5254954fb8611c0d67

Observation 5c6626de-6fc7-4472-9abd-8611b34178a7 · outbound

This paper cites On the design and analysis of llm-based algorithms.arXiv preprint arXiv:2407.14788, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs On the design and analysis of llm-based algorithms.arXiv preprint arXiv:2407.14788, 2024

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:10:22.199410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.613696Z digest=sha256:2104f6bdcf484ae3070681e9af638ce8f29525fa84b58f1d0142054163f77645

Observation 5c3a3ed5-4562-4bb9-a41d-72eecfe0a441 · outbound

This paper cites Grok, gemini, chatgpt and deepseek: Com- parison and applications in conversational artificial intelligence.INTELIGENCIA ARTIFICIAL, 2(1), 2025.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Grok, gemini, chatgpt and deepseek: Com- parison and applications in conversational artificial intelligence.INTELIGENCIA ARTIFICIAL, 2(1), 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.679570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.624596Z digest=sha256:c7165719989686cae92cdbecbbee5720936a32ed167b6e58361b3297d9593f3c

Observation ddd7c2a3-4a07-4ef7-8503-14a8219b0570 · outbound

This paper cites K2: A foundation language model for geoscience knowledge understanding and utilization.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs K2: A foundation language model for geoscience knowledge understanding and utilization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.653051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.633444Z digest=sha256:474cbab1c998c8f7e807d69d7d38e2bc571fba06d00551800766c68a5f1c1c7e

Observation 3a77f33f-629b-42ae-aa1d-af2ec3f14e6e · outbound

This paper cites A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data.IEEE Access, 9:165252–165261, 2021.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data.IEEE Access, 9:165252–165261, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.629090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.642177Z digest=sha256:10e1ac000c5dd783c0fa167f7d6203b6f57fcbb3a3b10b32d919a0ba251b0b3b

Observation 8e52f789-699d-415e-8981-809fe1d85431 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.649206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.649206Z digest=sha256:810c9ca89eb2ae8708fad3f76805c9720e7adaf0eca126cb0b751ed89b156dfd

Observation 901f78dc-51e1-4aa5-89ae-83a1174aeb2f · outbound

This paper cites The impact factor.Current contents, 25(20):3–7, 1994.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The impact factor.Current contents, 25(20):3–7, 1994

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.599611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.656055Z digest=sha256:17392e0f24a89445b63996505b20e6c95718f67315841a2393ccbec5c0daf5cb

Observation 558f8956-d940-4cbb-ac35-562400adbf7a · outbound

This paper cites The Llama 3 Herd of Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.662840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.662840Z digest=sha256:18faa409a0678316baab94ba8e0a95c7d4e09c2f909ab9d26b33131297690d1b

Observation a51b57c8-8b18-4140-ba82-5509949b35fe · outbound

This paper cites Llm-based code generation method for golang compiler testing.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Llm-based code generation method for golang compiler testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.566077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.674468Z digest=sha256:c7e15ec8660a0b8f49278afa0403bf90137591cafb4e01e7c867d81a91c0ddbc

Observation 95be0cb6-704e-402e-886b-af5bf384528f · outbound

This paper cites OpenDataLab: Empowering General Artificial Intelligence with Open Datasets.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OpenDataLab: Empowering General Artificial Intelligence with Open Datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.683625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.683625Z digest=sha256:3e2d12c49dc30863ee5e1967ce2cb97519d6af973544c00dee997e0908167489

Observation 91fc2325-9511-40b3-a819-e546cb97576a · outbound

This paper cites The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:10:22.028408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.695723Z digest=sha256:94c8e19dbebfeedababab8ede347028ad535c4eca6698f60497c620fd307b189

Observation b1fc3332-5753-4732-b197-51241d8d9f52 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.704332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.704332Z digest=sha256:60e96016af1c5b060790868f0245747b55cac58a16bbe4a63207fa59d4d10ff3

Observation 78b29462-ae4a-485d-b659-4a848770e410 · outbound

This paper cites The era5 global reanalysis.Quarterly journal of the royal meteorological society, 146(730):1999–2049, 2020.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The era5 global reanalysis.Quarterly journal of the royal meteorological society, 146(730):1999–2049, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.713964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.713964Z digest=sha256:6943f1c894ee365f5adf87b47a8888e4473f7d235b9970bedae2b9acd610dea5

Observation 281212e1-4636-41b7-acc7-a9c31a696dea · outbound

This paper cites Gpt-4o: The cutting-edge advancement in multimodal llm.Authorea Preprints, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpt-4o: The cutting-edge advancement in multimodal llm.Authorea Preprints, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.723245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.723245Z digest=sha256:55116600c6c36724189ccd65e97a17130ce96627e7cbb979381d3349cb539725

Observation ebe8e811-81a1-49bf-b046-a2957d4f76d5 · outbound

This paper cites Enhancing Large Language Models with Climate Resources.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Enhancing Large Language Models with Climate Resources

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.734284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.734284Z digest=sha256:23ab706c95f60b00e99e51b3e4768dd5e878e7f14c01c13057672d0c15adb636

Observation afffe1b5-8a9f-49a9-9371-5b9914af7231 · outbound

This paper cites an unresolved cited work.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:22.494299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.757242Z digest=sha256:08ef2532ee973ed47e9854cfea1f0b896b1906d0ca57f05547f6a480eadcd91e

Observation 74aae74f-2757-4bb5-b553-06560836c397 · outbound

This paper cites Iterative large language models evolution through self-critique.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Iterative large language models evolution through self-critique

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.471709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.766466Z digest=sha256:10f160246bccbf369a70ec08983e44a744591c7b4ef1f01a1fa45138eee9d5f6

Observation 6bc77e47-f0a3-4788-b6d7-8306b957569e · outbound

This paper cites LLM with Relation Classifier for Document-Level Relation Extraction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs LLM with Relation Classifier for Document-Level Relation Extraction

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:10:21.956033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.777262Z digest=sha256:eb9a69a16333f8c8448e2460d78dba5862719d68a9b694a8c5ca6904b74bef0d

Observation fe3cb26d-1bf5-4cb4-8269-3a815cbc21c2 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.798959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.798959Z digest=sha256:d908f66d18500edfa17ebfa2acda4cac19722f5ff1b2ce84659f2599b11d4cf5

Observation 872ab3e7-b821-49cc-87ec-146b1e0ff023 · outbound

This paper cites an unresolved cited work.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:22.445208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.819684Z digest=sha256:1f29a9edaa0384570a46da4d2a72ee8fa0c8dcd4028b69cba63dd5b8652bf3c6

Observation 056aff60-4819-4244-93b8-6950fc322455 · outbound

This paper cites DeepSeek-V3 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs DeepSeek-V3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.829089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.829089Z digest=sha256:405a8534b0a8a6b03a7f094dc5e06c6290432d4183c8639ac92966c6856927de

Observation 62bd0b92-6d7e-4daa-b372-7e9b269ae79d · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.840356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.840356Z digest=sha256:0bfef613a7999fbdd10fec30b56f5370c492909dd381840ce30c84e11d479362

Observation 6531b314-a0a7-4985-8138-d996dd5fdde2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.852395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.852395Z digest=sha256:42baf8778ba38ef55552479798b3c962ea684097cc55af84bcb82688a1eb8019

Observation 8f689b85-f820-48e3-a3e8-f959eef2b0ac · outbound

This paper cites The five environmental spheres.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The five environmental spheres

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.399166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.859905Z digest=sha256:17637a4b3471da7f16b2c76d9158b65edfc31dfc4dc6b760148d8304c96fb2e8

Observation bd80d183-00d0-4c2f-86be-20dacafbcab6 · outbound

This paper cites ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.868023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.868023Z digest=sha256:3a6a441b62b8f82fc44851ee5f6168d2586b9a23c126c50d8fb0282e9f7eecd8

Observation 45cf0f75-ab81-4122-ba23-88b942d26cd5 · outbound

This paper cites Seafloorai: A large-scale vision- language dataset for seafloor geological survey.Advances in Neural Information Processing Systems, 37:22107–22123, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Seafloorai: A large-scale vision- language dataset for seafloor geological survey.Advances in Neural Information Processing Systems, 37:22107–22123, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.378751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:10:20.876565Z digest=sha256:14b45427ad9b2c571e7408e3b82d0c3c10b17839ca1dae2c5f7d2187759136b1

Observation 4c1b245e-0f1d-4ba5-bac3-493010190ebf · outbound

This paper cites Is Temperature the Creativity Parameter of Large Language Models?.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Is Temperature the Creativity Parameter of Large Language Models?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.885364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.885364Z digest=sha256:16cc764f9775d8c27de1467b388509a06363e732185ddbc697076c3cf6cd56ef

Observation c50de15f-b858-4496-af4f-72d51508e97a · outbound

This paper cites Humanity's Last Exam.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Humanity's Last Exam

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.895287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.895287Z digest=sha256:496a0615be7deab75e4de66108b7d012931f577c20adfc5d2ff5513d48ccef4b

Observation 13166c0e-916b-40bc-b5bc-c9685bceabd0 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpqa: A graduate-level google-proof q&a benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.903533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.903533Z digest=sha256:3945c871b23d998527bc288f857b618f4cd6ac27e71403a4f03cbb9eaf10d644

Observation 7d5f817d-5d5f-484e-9271-46d7e81fbf16 · outbound

This paper cites From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.913666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.913666Z digest=sha256:a345013cb118a3167b66d7db6ea97128aefb6e6a7031bd444fa7d6b1ba7a0e08

Observation 2ea61cff-348e-4ac5-8f3c-92b75c277f79 · outbound

This paper cites Galactica: A Large Language Model for Science.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Galactica: A Large Language Model for Science

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.923953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.923953Z digest=sha256:4fe7f01a83173baa3c0e91a40327215a502587e1ba96ad21275c49caa5a1534e

Observation 673f422b-1eae-4dc7-b4f9-06cf2fd9c73e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.951647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.951647Z digest=sha256:95197ec6f08fe1265cb1ca48df66130fc47a57d6c2b5733c71b2e8be1afdf529

Observation bf4b7597-e101-4f0e-a4d5-bdccec3f0e9d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.957403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.957403Z digest=sha256:696d89ff6ef77a7f68834cad25c93dda30fa98c8295af8480c987c99529edc2e

Observation acd440be-6d2c-4734-9ae5-b41998e399ba · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.967014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.967014Z digest=sha256:09075baddf066559426fddbaafbfa3bd6a767aa33b502a4582594d4aae83187c

Observation 8756a6bd-781f-4f4d-9f0a-649e2e411800 · outbound

This paper cites ClimaText: A Dataset for Climate Change Topic Detection.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaText: A Dataset for Climate Change Topic Detection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.976823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.976823Z digest=sha256:3271fe02dacac42ada60d9f1c4bceb5cb04ba4a9545312604c47922529260dba

Observation 357c6f56-b065-4774-afe6-059089e7894a · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.984124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.984124Z digest=sha256:2e7a361362b1936f99547c2bd8cf61855f93b08e787b2823e51481e250b26924

Observation 12c2eb19-6730-442c-bc1a-4f20aabc21d0 · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.991495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.991495Z digest=sha256:e519cfc243295192439b98e7ebd92e858236f5b1f709ce71215b418b6a162b54

Observation 28c22c42-0dcb-4e72-a4ef-d3c17236d7a5 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.000707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.000707Z digest=sha256:fda6a906edc01921e65eeee366c40c7fc3c3c70612b0885205c66dde2c6ed8e3

Observation aa62ad69-0af6-4ff4-ba72-3673f29240f7 · outbound

This paper cites ClimateBert: A Pretrained Language Model for Climate-Related Text.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimateBert: A Pretrained Language Model for Climate-Related Text

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.012346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.012346Z digest=sha256:503bc7a61b07b02436e22523b1781504d647af33e5a982ac0275b9fb46779ff0

Observation ee202403-8be4-410d-aa55-a0fe69b11eba · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.021049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.021049Z digest=sha256:d4f1692f1d45f90cdcf3272db123ffe46e95de20cc55c9d3dbe256c8a8cefdec

Observation 882da276-877c-4230-8329-583e317e64c8 · outbound

This paper cites Measuring and Reducing LLM Hallucination without Gold-Standard Answers.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring and Reducing LLM Hallucination without Gold-Standard Answers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.033271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.033271Z digest=sha256:c0755f7db24ddf012e7d3a71532c2ab72c0fff80f670ec66d59437617433c9b7

Observation a3fa1203-58fb-4250-8f7a-2351f2c302e5 · outbound

This paper cites Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.043894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.043894Z digest=sha256:e7d90ba4799dc1788cdd4e2c9b37cfc62433386b7314bbd7374997e5355ea9bc

Observation 0a166818-bfdd-4da6-9091-d13638846a20 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.055953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.055953Z digest=sha256:83acd86cbf516b10de98bcf717484500d21828a796d34cdcca34993e6fd371f6

Observation b950559e-db08-45cb-9739-4e1e6c3a1833 · outbound

This paper cites Qwen2.5 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.069062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.069062Z digest=sha256:33fad1269de3906a7381d205d961f5c249d2ff80b65eae87bdb8609b6b8dc0a3

Observation 066a04ea-38ca-43f4-b377-dab3c81ac489 · outbound

This paper cites Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses.arXiv preprint arXiv:2410.07076, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses.arXiv preprint arXiv:2410.07076, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.082662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.082662Z digest=sha256:26ee8f4081d2e14b087b452908c35011e94be30d99ab660517a7dd6e4e106849

Observation f19a0ba4-1cee-4e09-b7f4-adc0898306ea · outbound

This paper cites EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.090064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.090064Z digest=sha256:6f3bcc0f6123c2814284c6628aa8e7611cf552d857dde67fdb95619a0323400e

Observation e5745806-da5d-4446-a65c-17f1f2f9bb1a · outbound

This paper cites ChemLLM: A Chemical Large Language Model.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ChemLLM: A Chemical Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.095513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.095513Z digest=sha256:726996b0fa958dc521941058c54babbfd0114d6136d7a878e6756847404de16c

Observation 315d0c97-3918-4b71-8097-c69c1a985ad6 · outbound

This paper cites Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.101708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.101708Z digest=sha256:87c5d7de592b03ff9d555dcc7c811ac62c6081b4a63d4a7959c5a8e14d649757

Observation de14b375-f8d0-4057-97a3-1c5fa38b8394 · outbound

This paper cites GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.111443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.111443Z digest=sha256:929fad2a46cb12d0c03626e5553f88495c5c9e1df1f8c1ae1793691f40dfcc69

Pith citing papers

Observation e1b52b5b-524f-4269-898e-da894a8fef87 · inbound

A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis cites this paper.

A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:22:26.646516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:22:23.855247Z digest=sha256:911555af3e6dd25c49cc87cc9ca7bc8aadcbc6cfa0be06a9e57f3916debb5927