Pith. sign in

Paper Citation Record · LEDGER

ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2410.05080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05080 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:22:36.386473Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ea070a9b-bd37-49f5-9c32-6d0352d86f8a · inbound

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model cites this paper.

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:14:04.994413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:14:04.994413Z digest=sha256:319a6ca69324aac4241f984b5232376724fbb20c62503fd0551faad4611e958d

Observation 29434ba2-8431-4776-9788-b08b6d2a65f3 · inbound

AIGS: Generating Science from AI-Powered Automated Falsification cites this paper.

AIGS: Generating Science from AI-Powered Automated Falsification ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:19.083819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:19.083819Z digest=sha256:8ca48cc915b0fc4c73189444fddb208bf91f3b73a4eadb262f9ec98d7c02f924

Observation cfb50cf9-b706-4218-be20-4918cdd4f817 · inbound

LLM4SR: A Survey on Large Language Models for Scientific Research cites this paper.

LLM4SR: A Survey on Large Language Models for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:39:25.157389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:39:25.157389Z digest=sha256:50681b1418c979386b3fd7d2a2062bafc9d91ed88b586d59e504eaef8402b2c3

Observation 56403468-0208-471c-8d4d-23555eee82b9 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.948145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:9c3685fc59f8de9336ab796546c6ce0316108fddaafc3527d90eaec4ebbfa197

Observation 54e66ece-9b82-4ad2-94d5-dfa5d0a6c5ed · inbound

Sparks of Science: Hypothesis Generation Using Structured Paper Data cites this paper.

Sparks of Science: Hypothesis Generation Using Structured Paper Data ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:22:36.386473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:22:36.386473Z digest=sha256:4e41a37a48d54673b7c63af74946c56505d466c145ee057671989ec85d0cdce8

Observation 35d435ba-a4ea-4159-9de9-391ecf6e661d · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.399364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:18e88821dde5649e45bb3e75aa79153f1895d097c94939ed58645033c46d04c9

Observation ef1c04c4-3289-40a4-9644-8b12fc8bbd36 · inbound

Can AI Agents Design and Implement Drug Discovery Pipelines? cites this paper.

Can AI Agents Design and Implement Drug Discovery Pipelines? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:44:03.815973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:44:03.815973Z digest=sha256:522010b9b72f19b4f397eaba831a2eee032a338e335651aef81f1affd59241ff

Observation 817f76da-d195-4d65-8fae-2f5f5bd6b4e5 · inbound

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies cites this paper.

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:26.078433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:53:26.078433Z digest=sha256:3faef1ac4456697700e26c3db658f2a240a75a95f24c89a5f4becd0f25860eb9

Observation 57b8442d-9b0d-43ef-8a8a-b526983ee376 · inbound

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research cites this paper.

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:32.267308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:32.267308Z digest=sha256:18b3feb122f3dc0642b8f8d84e1ce82204b2c7b400ad2263be9b93b3595a1bf7

Observation 88508b9d-2c41-4609-a8b3-28574d48c1da · inbound

EXP-Bench: Can AI Conduct AI Research Experiments? cites this paper.

EXP-Bench: Can AI Conduct AI Research Experiments? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:20:42.230426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:20:42.230426Z digest=sha256:6ff1d6bdadce1705e774c99fae791fc4b5cf56cfae0be483d436d333f8503e01

Observation b604650e-6e04-4d51-8e11-39f3f68716b7 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:50.860148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:50.860148Z digest=sha256:a9c093dcc1702da376e45fa12fa88164268c02b66a24cdb9563274506517c3b7

Observation d9e984a3-57ee-4d7a-bf01-87859956317b · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.940291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.940291Z digest=sha256:35d99826ec1aa21d9b143f32202212bdc22143f05f73e377011870fd57fcabb6

Observation 6a271e68-1308-4909-8500-8452e86d43b9 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.467715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.467715Z digest=sha256:36e173e5a555643f54f82ecf35398262969f09eb1a9863f88c2b6c27d26211a3

Observation 8c408575-b0e8-4906-a807-f5ffc28cd7db · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:07:39.431114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:ef7071aa3ee774dd9b93625e330c6fff4dfacc057f74d9d379ba52a2d0379383

Observation 0b835d3e-c055-4cc5-a265-10a43a846660 · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.881460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.881460Z digest=sha256:ff50860cbbf8e139619ad8b2f7840b710d2540c973a3499612b27777ff5e5e3c

Observation 5d1f6fcf-7734-4252-98b0-c18c8d5ba657 · inbound

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents cites this paper.

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:52.985789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:52.985789Z digest=sha256:2b35405ca2b1c9c0a1aab6ac4852570df1c200cd6b3f833753959b20da7f880c

Observation 3e509b4e-623e-4f89-aee5-3cf5d378c250 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.564539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.564539Z digest=sha256:74b892ee5c908ffee3f9667e8429b105cc73cf3e65d87557b99e46b15280e0ff

Observation 43a63161-5ccf-4255-a97f-5f0e215d83d2 · inbound

Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation cites this paper.

Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:05:06.557517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:05:06.557517Z digest=sha256:f3c7b48385a7e930a567effdd768f9b9867b372f165542280171eb6abed1bc76

Observation cb1f5943-ef43-4962-834f-af9a07824c62 · inbound

RExBench: Can coding agents autonomously implement AI research extensions? cites this paper.

RExBench: Can coding agents autonomously implement AI research extensions? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:37:08.883664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T07:33:39.675929Z digest=sha256:d838417c8cdb11dbd945077d6d655b765f665bbb1b240976a4c858912477081e

Observation 23eb0a43-0b73-46a5-a1dd-d05265413501 · inbound

AI4Research: A Survey of Artificial Intelligence for Scientific Research cites this paper.

AI4Research: A Survey of Artificial Intelligence for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:12.313531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:12.313531Z digest=sha256:5c84c9f5ff62793f71d3ce2d6debd9f863ab3490b2475673886dda5284325d80

Observation 41d00269-018b-4580-a70c-f1251c57e29f · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.664362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:f8f1118dd396fa133585b9fbb9d8dc458b9c3b40ff4f0b6dbbc9840905441696

Observation 2094d729-a247-4398-969c-1bfd1da0a94a · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.520560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.520560Z digest=sha256:d7b7d68ab9a7ee754a2b6fc0e428fb97eee15dbc97718fac8c3ea93277308bdf

Observation 91b010aa-dd33-4aed-8976-093534736d28 · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:14.623103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:14.623103Z digest=sha256:7a32a1b4328c877081ab12a0f101c7bd0db2e2761f2f50e0020d4f99c880da56

Observation 4ebaa39f-9807-4cd0-99e4-a865f665e472 · inbound

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation cites this paper.

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:24:25.055333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:24:25.055333Z digest=sha256:9358d2f713f094dd994e65654b7ebdb1a7fb58a69301190fc337f7286fa536e4

Observation 56d6a0f3-0ff2-47df-9368-50f9ee986c56 · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:06:32.002005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:a70d065b73398cbe50c05ac95fb1175223606c72474711e238721220638cbcf4

Observation 6d279899-50c3-4d4f-9689-283aec17daf3 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:36.771781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:36.771781Z digest=sha256:d25eaf13f582a1c33c7caaf83f6c2c99e3b2a60fcca36b8a9c557934a803736a

Observation 808a1a95-e296-4a59-b58f-27eb79084f56 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.120470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:09a8d8db3f826ae742ff21feb2116b9b210d870366693ccaea0d31402f881a0c

Observation 23341957-a990-4a2c-91e3-80cffd3963ab · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:11.594053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:11.594053Z digest=sha256:b16ef565559af45fe5776f6e34502696a2bb7cba7d7972ff92ec3575ef044208

Observation 3bf352c4-3f33-44a0-9224-6a15da628d5a · inbound

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence cites this paper.

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T21:00:05.911327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:00:05.911327Z digest=sha256:c7cdb2de577c9ce70f541697282d16fea8f31241cc36a9ae6fa1a75f4689c875

Observation 88755ec6-6b79-4c94-94c8-2e215a67053b · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.990899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:c0af99c0dac55aea0233b044ff0d14c72a72bff8366377061fc8c845f140153f

Observation 512bc90b-06e6-4d78-9603-fd06ed6e08fb · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.389705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T05:59:01.010437Z digest=sha256:ad935e8df0adf46fe7b707ea76851c5ccb2a10391c3b42eccf59fa7c3cc6e840

Observation eb0aac0c-7f63-4c95-9f21-f7fb03cd1434 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.762793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:edc132399b178bb3f2779811267ddf19b3e6473cea766b8ee809b13f7feee462

Observation e3d50f7d-c52f-4211-b7dd-ab70da14f2fd · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:07.053196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:dd1be3ab65cf5ac68d20e23d1d4d31260521c660b7b3b784edcd44fa06befea9

Observation 008b456a-fafd-4dc0-bca1-ff4fbf8f2be5 · inbound

Agentic-imodels: Evolving agentic interpretability tools via autoresearch cites this paper.

Agentic-imodels: Evolving agentic interpretability tools via autoresearch ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.314009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T16:37:43.371592Z digest=sha256:30e3357d34e5b4ec9418cdd31ae2ea53d0de5cee0017af34c8aad49160fb866c

Observation 7db4a2ba-ff3c-4566-88de-01bd7160842b · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:31:07.935615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T17:28:41.217810Z digest=sha256:f27996c2d3b705d2153aeb3cd4d67a3782e6ee07e558ddc5895941a45c63b240

Observation d1d1eb3b-a264-41a7-9589-301092f1dc7f · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T00:13:52.856073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T00:13:07.546472Z digest=sha256:fa2f6f2601b978da4a0252b2e74359dedcebce785607cc2929a03b22087b98c0

Observation bd5541f9-db46-4f0e-9f26-4ed70b4ec7d4 · inbound

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy cites this paper.

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.419263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:14:48.091245Z digest=sha256:3480d602a2e230228b74153e4d4fdbf0b75e08c5f5ad331a68a7186090a256a3

Observation 7c24fc5b-fbc7-401a-8e85-214ee5d8e4b9 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:42:59.091537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:23:34.514071Z digest=sha256:53c0897da7b6f52157da38d11be5a7c349e26662ff8c907db1c37ead2340e4d1

Observation f9d3a9a8-70bf-4e06-8679-25a5b4d2b280 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:55:04.185995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T04:51:17.519200Z digest=sha256:34d20063ae164a82ba2e8f6ab7f1744f2ba8f8ea3b1f0468eab59c4d6fafb045

Observation d9150b61-a096-4425-8ef3-c022481c9690 · inbound

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents cites this paper.

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:49:48.155452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T05:45:13.652112Z digest=sha256:743d6f446fe24bc6a8f1aefda339a31d912793b0a99731893610fa01199fb08c

Observation 9ae8b600-47c5-42ad-9d29-ecefac5a3aa4 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.130932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:695d3fb25ea431f4fcf5959b58ae373ccdfb7061c60dc7c621b1b36e8f5f84d6

Observation 4d64616a-0ed4-4864-8077-eb79813f0a3a · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.416859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:5ffe7b06dcfdbc1751c6604e1e1e52b9d912963f10c3f999fb97c0064f5b0b93

Observation 6c4e6d4f-f2e8-47be-8365-17f4c87e9d6a · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.919428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:610735c16db271fd0f57b0139c108c3c963431d878dab9afbcad03735e87a6e2

Observation b816edfd-d501-41f4-a7d6-3f0852dfaea2 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.174291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:22d7b37493ed56a091fb8db6f05be5071a56995ddd8875d1656c67c9354e2178

Observation f05f4a13-68ad-4357-8163-5a251d1de504 · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:58:11.275989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:e093defc255f581c450a379c3c5d16752d75108ad30115b372120fe7564dfb23

Observation 64537fce-d2f5-4113-82fb-612ab3852546 · inbound

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research cites this paper.

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T12:12:07.893078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T12:08:10.552789Z digest=sha256:b3dac096470e46302258c7341e8d75ac35de5d207b461a42ba1aff5add593d53

Observation f4bb8ced-03cd-4618-af61-ea1daf4023b1 · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.268768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:7f04d5b20634881ace7b580aabfa13af11ec2fcff5fe9614f121dc218f088dd5

Observation eec0945f-3e54-43de-a0ad-010904ff15c6 · inbound

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters cites this paper.

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:32.158536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:14:26.278017Z digest=sha256:b636ddf5ae79e6eef4bfc758db438c89c1ed96bf373fac47e7b7faed5b778515

Observation 2584bc5b-0ca1-4b18-8b8b-281f88b05362 · inbound

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters cites this paper.

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:23:33.009993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T05:25:29.078764Z digest=sha256:22ee56bc19a4337e13cc14c3baeba3a0184604585a30d48354989a6e8bd17e99

Observation 56fd1037-0225-4f58-8cd5-8a890818a602 · inbound

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories cites this paper.

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.774873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T13:43:02.919248Z digest=sha256:c7b6538274906b834104498ba3502b26017c206816540492e40306e1231f8ece

Observation de9f9559-3f59-4240-9449-302bfad1419c · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:40:47.024579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:72592b523b78819133da7c4f5dd546a23bf6563316afcbc314eb2f00ea74464b

Observation 5e743964-8983-4ece-a96c-7cb1ab00d7b8 · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:44.063091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:a7d084f4a063238f7d0fb05d2eba82d279c2febda83264265744c4b8e3cb4759

Observation bd4861ec-7298-450a-83f5-50c4d597d14e · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:48:45.699996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T03:48:51.497061Z digest=sha256:886b10ba81fd7b8c2bdf705bae4f8ca07e175728733af0e548961c634782d359

Observation 7a4b9f53-6523-4f29-94fa-10146fbcdeeb · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:54:38.905083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T10:23:03.764108Z digest=sha256:62ab54d449e4448ceaf76203639ec835b1bb760456f1d7aa41e4a39a5717488e

Observation cc54a9da-76af-46ac-abc7-ab3a7a9f9119 · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.722606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T22:05:37.948336Z digest=sha256:95c45d18881371ac370b08438544905259b3a4bc20227fe83999fd245d32fe0f

Observation 20aa0762-abaa-4548-a728-44d9794ff469 · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T11:08:45.071671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:08:45.071671Z digest=sha256:b4baf293232e86a3498add4a86832c2e6403c178e97d97014cd7d401abc74934

Observation ffaa95bd-4fa8-4215-a572-169e113bc2a2 · inbound

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation cites this paper.

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.358749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:07:49.603969Z digest=sha256:da1e1030370fbfd6c60b83d279acc4a868628ed5964c464d3385d4ad1813736f

Observation 97ae98a8-5aca-4e41-8377-b824b5aae258 · inbound

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents cites this paper.

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:18:59.717212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T23:40:53.439549Z digest=sha256:a6b5d3f654bd61128b99db9e557ee82577e8368e049f737e87ebf49c4c91288d

Observation 4ebafdfb-d718-4a22-906d-22d78ff414fc · inbound

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark cites this paper.

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.272487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T19:17:17.373463Z digest=sha256:ab461a9762309389fcf6fc0a124606a0dcd675af7cf39a017607511ca0280439

Observation 44d78091-2566-498b-b749-93971a525d1a · inbound

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution cites this paper.

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:39:30.801028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T17:45:49.272070Z digest=sha256:564d41d1414ba1e2fc8beb1b3aed0fc625f2239a8cfff606b65f771c0373e661

Observation 39f1b902-6d5b-4f37-bece-6c0a835556e3 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.683593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:84900dee8a65c09d78484c700566c20039c81ad914227e04908ada0208c7c85c

Observation 6aca5516-39b5-4cb5-a91b-53c26b4307a6 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 188

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:45:55.719982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:f6ffe1824c3d3228366452739e9ce932e0fda60343ef060953a136c0c9deae0e

Observation ebd57547-9138-42c3-9bcf-b982a4fac180 · inbound

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems cites this paper.

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T00:55:17.107245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:55:17.107245Z digest=sha256:f1a3ddb3eb3e2c358ea2f0c405f00cf04943af885766dd97ff3f004196e8b21e

Observation f3b9591a-9c4c-4707-b2a8-8863478fcf5e · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:806ec7ca921107a410f8d5a21d8d85ea39775f975bb371047ebc5c7ded8d50b7

Observation 8c12ae4c-2cc0-45c9-bd76-b4e32ef82e35 · inbound

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation cites this paper.

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T13:28:32.903465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:28:32.903465Z digest=sha256:c26eb68795f95705e2dc8ac0fa47f67a1905e8901bbe1b5faf8ae758231cf694

Observation 487e51cc-4ac0-47cd-a22e-49062cd04b05 · inbound

SciDataSailor: Deep Scientific Data Exploring cites this paper.

SciDataSailor: Deep Scientific Data Exploring ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:02:37.896688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:02:37.896688Z digest=sha256:22d8bfe8fb4ccfea6287af82ec2a23ea0f842e5b89f203aeccdf5facfe0d97bb

Observation a3ab6b94-0a2d-4abd-b2ed-7c5c83063fe8 · inbound

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities cites this paper.

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:49:30.068458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:49:30.068458Z digest=sha256:f4e76f6ba8c8879aaeabbbe104078f2add614e7a1babc9d02e890238736772d4