Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2410.05080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:22:36.386473Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
6
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation ea070a9b-bd37-49f5-9c32-6d0352d86f8a · inbound
AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29434ba2-8431-4776-9788-b08b6d2a65f3 · inbound
AIGS: Generating Science from AI-Powered Automated Falsification ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb50cf9-b706-4218-be20-4918cdd4f817 · inbound
LLM4SR: A Survey on Large Language Models for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56403468-0208-471c-8d4d-23555eee82b9 · inbound
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54e66ece-9b82-4ad2-94d5-dfa5d0a6c5ed · inbound
Sparks of Science: Hypothesis Generation Using Structured Paper Data ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d435ba-a4ea-4159-9de9-391ecf6e661d · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ef1c04c4-3289-40a4-9644-8b12fc8bbd36 · inbound
Can AI Agents Design and Implement Drug Discovery Pipelines? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 817f76da-d195-4d65-8fae-2f5f5bd6b4e5 · inbound
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b8442d-9b0d-43ef-8a8a-b526983ee376 · inbound
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88508b9d-2c41-4609-a8b3-28574d48c1da · inbound
EXP-Bench: Can AI Conduct AI Research Experiments? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b604650e-6e04-4d51-8e11-39f3f68716b7 · inbound
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e984a3-57ee-4d7a-bf01-87859956317b · inbound
TextAtari: 100K Frames Game Playing with Language Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a271e68-1308-4909-8500-8452e86d43b9 · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c408575-b0e8-4906-a807-f5ffc28cd7db · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0b835d3e-c055-4cc5-a265-10a43a846660 · inbound
Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1f6fcf-7734-4252-98b0-c18c8d5ba657 · inbound
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e509b4e-623e-4f89-aee5-3cf5d378c250 · inbound
Deep Research Agents: A Systematic Examination And Roadmap ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a63161-5ccf-4255-a97f-5f0e215d83d2 · inbound
Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1f5943-ef43-4962-834f-af9a07824c62 · inbound
RExBench: Can coding agents autonomously implement AI research extensions? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23eb0a43-0b73-46a5-a1dd-d05265413501 · inbound
AI4Research: A Survey of Artificial Intelligence for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d00269-018b-4580-a70c-f1251c57e29f · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2094d729-a247-4398-969c-1bfd1da0a94a · inbound
Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b010aa-dd33-4aed-8976-093534736d28 · inbound
How Far Are AI Scientists from Changing the World? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebaa39f-9807-4cd0-99e4-a865f665e472 · inbound
GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d6a0f3-0ff2-47df-9368-50f9ee986c56 · inbound
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d279899-50c3-4d4f-9689-283aec17daf3 · inbound
How can we assess human-agent interactions? Case studies in software agent design ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808a1a95-e296-4a59-b58f-27eb79084f56 · inbound
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23341957-a990-4a2c-91e3-80cffd3963ab · inbound
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf352c4-3f33-44a0-9224-6a15da628d5a · inbound
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88755ec6-6b79-4c94-94c8-2e215a67053b · inbound
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 512bc90b-06e6-4d78-9603-fd06ed6e08fb · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eb0aac0c-7f63-4c95-9f21-f7fb03cd1434 · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e3d50f7d-c52f-4211-b7dd-ab70da14f2fd · inbound
AI scientists produce results without reasoning scientifically ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 008b456a-fafd-4dc0-bca1-ff4fbf8f2be5 · inbound
Agentic-imodels: Evolving agentic interpretability tools via autoresearch ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7db4a2ba-ff3c-4566-88de-01bd7160842b · inbound
Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1d1eb3b-a264-41a7-9589-301092f1dc7f · inbound
Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bd5541f9-db46-4f0e-9f26-4ed70b4ec7d4 · inbound
gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7c24fc5b-fbc7-401a-8e85-214ee5d8e4b9 · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f9d3a9a8-70bf-4e06-8679-25a5b4d2b280 · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d9150b61-a096-4425-8ef3-c022481c9690 · inbound
Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ae8b600-47c5-42ad-9d29-ecefac5a3aa4 · inbound
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4d64616a-0ed4-4864-8077-eb79813f0a3a · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c4e6d4f-f2e8-47be-8365-17f4c87e9d6a · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b816edfd-d501-41f4-a7d6-3f0852dfaea2 · inbound
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f05f4a13-68ad-4357-8163-5a251d1de504 · inbound
How Far Are We From True Auto-Research? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64537fce-d2f5-4113-82fb-612ab3852546 · inbound
Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4bb8ced-03cd-4618-af61-ea1daf4023b1 · inbound
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eec0945f-3e54-43de-a0ad-010904ff15c6 · inbound
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2584bc5b-0ca1-4b18-8b8b-281f88b05362 · inbound
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 56fd1037-0225-4f58-8cd5-8a890818a602 · inbound
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation de9f9559-3f59-4240-9449-302bfad1419c · inbound
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5e743964-8983-4ece-a96c-7cb1ab00d7b8 · inbound
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bd4861ec-7298-450a-83f5-50c4d597d14e · inbound
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a4b9f53-6523-4f29-94fa-10146fbcdeeb · inbound
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc54a9da-76af-46ac-abc7-ab3a7a9f9119 · inbound
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 20aa0762-abaa-4548-a728-44d9794ff469 · inbound
MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffaa95bd-4fa8-4215-a572-169e113bc2a2 · inbound
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 97ae98a8-5aca-4e41-8377-b824b5aae258 · inbound
SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ebafdfb-d718-4a22-906d-22d78ff414fc · inbound
Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 44d78091-2566-498b-b749-93971a525d1a · inbound
Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 39f1b902-6d5b-4f37-bece-6c0a835556e3 · inbound
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6aca5516-39b5-4cb5-a91b-53c26b4307a6 · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 188
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ebd57547-9138-42c3-9bcf-b982a4fac180 · inbound
Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b9591a-9c4c-4707-b2a8-8863478fcf5e · inbound
Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c12ae4c-2cc0-45c9-bd76-b4e32ef82e35 · inbound
SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487e51cc-4ac0-47cd-a22e-49062cd04b05 · inbound
SciDataSailor: Deep Scientific Data Exploring ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ab6b94-0a2d-4abd-b2ed-7c5c83063fe8 · inbound
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.