Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:21.111443Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2505.17139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:21.111443Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:22:23.855247Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:22:26.516274Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 12e6ab02-ab96-4435-9f59-007f7f320169 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7f9f8a-0dc7-4e7c-b5ab-106c5b6bd07a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OceanGPT: A Large Language Model for Ocean Science Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9a4a9b-84d1-49dc-8890-d6f090abcec9 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Matplotlib and seaborn
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee12df61-3140-427d-89bd-3baa7c2563f7 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs This reference does not exist: an exploration of llm citation accuracy and relevance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 03694014-52e2-4d23-a59a-ca3afde7f50d · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6626de-6fc7-4472-9abd-8611b34178a7 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs On the design and analysis of llm-based algorithms.arXiv preprint arXiv:2407.14788, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c3a3ed5-4562-4bb9-a41d-72eecfe0a441 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Grok, gemini, chatgpt and deepseek: Com- parison and applications in conversational artificial intelligence.INTELIGENCIA ARTIFICIAL, 2(1), 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ddd7c2a3-4a07-4ef7-8503-14a8219b0570 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs K2: A foundation language model for geoscience knowledge understanding and utilization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a77f33f-629b-42ae-aa1d-af2ec3f14e6e · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data.IEEE Access, 9:165252–165261, 2021
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e52f789-699d-415e-8981-809fe1d85431 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901f78dc-51e1-4aa5-89ae-83a1174aeb2f · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The impact factor.Current contents, 25(20):3–7, 1994
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 558f8956-d940-4cbb-ac35-562400adbf7a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51b57c8-8b18-4140-ba82-5509949b35fe · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Llm-based code generation method for golang compiler testing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95be0cb6-704e-402e-886b-af5bf384528f · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OpenDataLab: Empowering General Artificial Intelligence with Open Datasets
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fc2325-9511-40b3-a819-e546cb97576a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1fc3332-5753-4732-b197-51241d8d9f52 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b29462-ae4a-485d-b659-4a848770e410 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The era5 global reanalysis.Quarterly journal of the royal meteorological society, 146(730):1999–2049, 2020
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281212e1-4636-41b7-acc7-a9c31a696dea · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpt-4o: The cutting-edge advancement in multimodal llm.Authorea Preprints, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe8e811-81a1-49bf-b046-a2957d4f76d5 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Enhancing Large Language Models with Climate Resources
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afffe1b5-8a9f-49a9-9371-5b9914af7231 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74aae74f-2757-4bb5-b553-06560836c397 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Iterative large language models evolution through self-critique
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bc77e47-f0a3-4788-b6d7-8306b957569e · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs LLM with Relation Classifier for Document-Level Relation Extraction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe3cb26d-1bf5-4cb4-8269-3a815cbc21c2 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872ab3e7-b821-49cc-87ec-146b1e0ff023 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 056aff60-4819-4244-93b8-6950fc322455 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs DeepSeek-V3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62bd0b92-6d7e-4daa-b372-7e9b269ae79d · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6531b314-a0a7-4985-8138-d996dd5fdde2 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f689b85-f820-48e3-a3e8-f959eef2b0ac · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The five environmental spheres
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd80d183-00d0-4c2f-86be-20dacafbcab6 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cf0f75-ab81-4122-ba23-88b942d26cd5 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Seafloorai: A large-scale vision- language dataset for seafloor geological survey.Advances in Neural Information Processing Systems, 37:22107–22123, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c1b245e-0f1d-4ba5-bac3-493010190ebf · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Is Temperature the Creativity Parameter of Large Language Models?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50de15f-b858-4496-af4f-72d51508e97a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Humanity's Last Exam
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13166c0e-916b-40bc-b5bc-c9685bceabd0 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpqa: A graduate-level google-proof q&a benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5f817d-5d5f-484e-9271-46d7e81fbf16 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea61cff-348e-4ac5-8f3c-92b75c277f79 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Galactica: A Large Language Model for Science
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673f422b-1eae-4dc7-b4f9-06cf2fd9c73e · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4b7597-e101-4f0e-a4d5-bdccec3f0e9d · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd440be-6d2c-4734-9ae5-b41998e399ba · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8756a6bd-781f-4f4d-9f0a-649e2e411800 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaText: A Dataset for Climate Change Topic Detection
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357c6f56-b065-4774-afe6-059089e7894a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs MinerU: An Open-Source Solution for Precise Document Content Extraction
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c2eb19-6730-442c-bc1a-4f20aabc21d0 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c22c42-0dcb-4e72-a4ef-d3c17236d7a5 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa62ad69-0af6-4ff4-ba72-3673f29240f7 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimateBert: A Pretrained Language Model for Climate-Related Text
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee202403-8be4-410d-aa55-a0fe69b11eba · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882da276-877c-4230-8329-583e317e64c8 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring and Reducing LLM Hallucination without Gold-Standard Answers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fa1203-58fb-4250-8f7a-2351f2c302e5 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a166818-bfdd-4da6-9091-d13638846a20 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b950559e-db08-45cb-9739-4e1e6c3a1833 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Qwen2.5 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066a04ea-38ca-43f4-b377-dab3c81ac489 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses.arXiv preprint arXiv:2410.07076, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19a0ba4-1cee-4e09-b7f4-adc0898306ea · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5745806-da5d-4446-a65c-17f1f2f9bb1a · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ChemLLM: A Chemical Large Language Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315d0c97-3918-4b71-8097-c69c1a985ad6 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de14b375-f8d0-4057-97a3-1c5fa38b8394 · outbound
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b52b5b-524f-4269-898e-da894a8fef87 · inbound
A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.