Pith. sign in

Paper Citation Record · LEDGER

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

As of 31 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2506.02314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02314 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T14:26:52.414276Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:23:24.671544Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8845b873-a0d6-4b73-a5a1-001b7f67e920 · inbound

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents cites this paper.

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:15:28.008514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-21T18:15:13.924995Z digest=sha256:add4c12475ffaa1f7b05936b6e505799cac284fbb46539c526d260616b29f6ec

Observation 4ad5c576-5f75-422e-89b7-6473640b0e45 · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:06.935796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:3bf1bcd266c22fac5d46dd7dedfeded27790b398ef4dc3c9fd8125bf837b91a2

Observation 7d50b9dd-91ab-491b-b104-8c950183688f · inbound

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification cites this paper.

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.271560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-13T17:32:06.864338Z digest=sha256:026dccbb50f080bdc13123a00dca4c4bc022d71d075479812d057e23334a338c

Observation 9827d15f-5293-4b57-aa18-9223ba4c8115 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.778992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:5d8a8ad25afd432c3b2def67b1525d11c7186660327e4900658b219399d3dcb6

Observation eabd1c0e-4c1b-4689-a89f-5ed82c8c11d7 · inbound

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories cites this paper.

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:01:04.244734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T02:30:39.346257Z digest=sha256:d79a41a945b48c000ce87c1d2e2f7c30f7b9519eb78cf9bd9f102585b0678c7b

Observation eddb464f-bb32-48c8-91c2-7cbca15f67af · inbound

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench cites this paper.

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:52:43.035333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-19T17:49:01.198956Z digest=sha256:4240a256e2c3cfe6c72e1e547c5b55909d8688d2e7f536bdb7435199fb7b72bf

Observation f78992f4-4ce3-462a-b6a9-68a132b93bb6 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.142766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:63e5f07f5d1d0ccd0c5d54fcd24b1398bdd402425914b532c62a66a1cf6a9ddb

Observation 78d3d367-84ca-450b-bab8-24e16520e1a2 · inbound

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents cites this paper.

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:57:53.363396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-19T23:54:15.987953Z digest=sha256:7d6dc3dc6c7626d0a1a6082cfb7a62fb92d522cb458fb8ac0f1daf8e340befa0

Observation 87403219-8c32-45e1-bfd6-113e90b5ff56 · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.684486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:e7dae4fdd65c983d7db8ccef492856140abc5268ac3a6c24ea7ed52d9503e8ef

Observation 943dad94-76be-4e79-8e7e-ac03b93f6b2a · inbound

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence cites this paper.

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.672894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-29T12:13:58.111299Z digest=sha256:dac46dcc80d61c76c669f9984c8ed49738ab78985fe10a6321637bda2a3b688c

Observation 88708047-8690-4df7-a06e-1ca8438e43fc · inbound

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis cites this paper.

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T14:26:52.414276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:26:52.414276Z digest=sha256:3010d0122862d0b1051ffc5cb2c5a60c4852d4eff7c7f88789e1d5280fc8fae6