Pith. sign in

Paper Citation Record · LEDGER

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2503.21248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21248 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:17:18.385882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a622c17-1c75-40b4-9ecb-f3448cf8d2a0 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 235

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:38.190904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:da9be1c9dc034ac93513aece98fdb6276550615d8d498e207a3987f6bbf94b13

Observation 5f454d90-f8c8-4364-9ffb-00b377deb460 · inbound

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research cites this paper.

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:01.581385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:48:44.217655Z digest=sha256:e8e267ccab973c7fb6d5805245b839a760d92f2fe23d900b184885b892eb040d

Observation d5f29e90-b949-46ec-9c5c-79d65cf60fb8 · inbound

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights cites this paper.

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T05:17:18.385882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:17:18.385882Z digest=sha256:6ae3b926c3a0751fd34760be422209c85c32c328e30f684f1a209a7496726862

Observation 87d68fee-abb4-4001-ae3f-6ece1f560ad4 · inbound

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery cites this paper.

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:46.930452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:46.930452Z digest=sha256:eeb868ef208f6186bf97bbab20542b8c8692f24582d915127f96cad0ad195bdf

Observation c836fe18-46e3-40a0-a902-a9541aab91c0 · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:16:06.789382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:7ea45765199db70fbebf82009655944cb83f35d7029fde091345f0954beb434e

Observation 6e5068c3-ca03-403c-95f1-78352ef8f733 · inbound

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification cites this paper.

AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:31:13.776475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T05:23:18.294600Z digest=sha256:21e3d2d754dbbf50318fe66b5c5d6091f6f6ec4b4151cc0d56b1fcca19f94982

Observation c412b56e-455c-434d-a41d-c14083931d31 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:25.167029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:0dbce72cebc26a7ba3c708d25813810b37a4814562c2d56c68fbb658b16bd09c

Observation a985dfca-1072-4802-a6a1-9063a0c4f531 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:46.055254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:c4a67c3d6b81d0a952d5a397b1da5709b41a6b0d76a60b0ac45c237cc595e2bf

Observation c7f875f8-afcb-4b52-b7bb-cccf4226276b · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:9d5f491f5bb0c071e37acf61a5148276a75b7dacf8b50f09d820c98298bc3c3d

Observation 221c4c7c-b92b-4040-b84f-f23b649768d4 · inbound

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design cites this paper.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.068813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:93c9d1c17b3554a55515def597eca81a235b4d1d6cc207bec47ff8c113884fdd

Observation 7cdf9a01-c210-4718-b773-a5d5abdedb0a · inbound

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI cites this paper.

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:21:07.442521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T05:18:15.358571Z digest=sha256:651ffc918f30553cc938bd7b2eec5a2bf883446719fcf2f0a9dcd982dac8d059

Observation 11219177-21d5-4887-b9c5-be5225a21935 · inbound

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI cites this paper.

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T13:31:44.953243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:31:44.953243Z digest=sha256:33440cccde56568f5ad15f77146aa4fe988a64ee32a36d414cd741630a06e40c

Observation fcbbcd39-7d85-45c5-90c6-17ffaa4d8414 · inbound

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery cites this paper.

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:50:21.922451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T04:46:43.679185Z digest=sha256:77f379fa007e5b318388cb3b70afa5c7180ff6e1c2e75faa38b21f8294033586

Observation f8b623b0-94f6-45fb-b711-c763fca21614 · inbound

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence cites this paper.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:23:59.088553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T21:19:03.281629Z digest=sha256:64b73cef761fb3a69e4a0b8a45c1f32f0b305e2515c28e3466b956a7e6fcd56c

Observation 772902d4-0768-428d-a06f-4b1ddf8fb818 · inbound

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research cites this paper.

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-28T12:12:07.898907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T12:08:10.552789Z digest=sha256:f0bf2cee3203ced1d33f0d0211e1df1af269d39562b53485ecf8af7cb95fbfe6

Observation ca7970c7-2ef8-4d1f-b82a-3c2d6b8fce71 · inbound

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations cites this paper.

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:26.421873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T19:00:15.920341Z digest=sha256:654e9aa285dc85570fa99a24e27be3a38946404338229366d3e33692c9a5e1e1

Observation 7465a901-1b6b-4988-b926-d283c48ba0f7 · inbound

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations cites this paper.

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:09:00.647335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:09:00.647335Z digest=sha256:2a7c29cf47faa31b51021309dce75288af00d3fe1c5ba1be06ac0dba6d4c5f74

Observation 7617b107-6444-487f-af8a-02423656c2e5 · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T11:28:04.228018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:69135dc6bbab14558a087c0415a1dc596f7a2def85fc9bd1e7c4f636a9c669d9

Observation ff822da0-cba6-43d1-ba64-c14e416634ec · inbound

Measuring the Gap Between Human and LLM Research Ideas cites this paper.

Measuring the Gap Between Human and LLM Research Ideas ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:26:56.105887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-02T12:22:54.748120Z digest=sha256:4b7b3b351f97e5f5b4ec1fb883cfb62904183c8d1d093b79d7b01f5696118bd3

Observation 9e5c07ea-8d73-4a72-ad22-c6ba4b08073a · inbound

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games cites this paper.

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T20:19:27.650696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:19:27.650696Z digest=sha256:b1629da390b3a74b5c6426267d101668fd3c4688e4274740daf384909000c5ae

Observation 5e3cc368-8abf-45fc-92aa-83f6af3c85c2 · inbound

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes cites this paper.

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T19:11:37.739920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:11:37.739920Z digest=sha256:d69c4c9466754e654c2aa68ecdac493d981d453e8c68295845231c19833160f9