Pith. sign in

Paper Citation Record · LEDGER

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

As of 19 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 17 inbound Pith citation observations for arXiv:2506.02314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02314 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:30:04.614811Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:46.242048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:23:24.671544Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 691111d9-f02b-4677-9914-91bc4770ad7d · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.291244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.158768Z digest=sha256:5301c111a3339327bf49c9947fd5846ed5e1ab1acac0fdc05a17e3e790cf42a3

Observation 2ddd7919-e2b9-479b-afd5-50c3228893df · outbound

This paper cites calculate area.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code calculate area

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:30:05.132937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.233185Z digest=sha256:588a9a323c41bb7ab260c192341f8820e7e781be9f3896c12031876345573668

Observation 66794533-d14b-4e1b-9c2e-d2ffd3814057 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.012101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.285489Z digest=sha256:5c3fc065457e57e56c6cfa524e0d9f59080cd0c0dfc11583f1407e78430aa3eb

Observation 25375523-4700-4625-b181-7f40447c8193 · outbound

This paper cites 19 Here is the code that you need to complete: {context_code_str + masked_code_str} Please implement the missing code in the TODO blocks.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code 19 Here is the code that you need to complete: {context_code_str + masked_code_str} Please implement the missing code in the TODO blocks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:30:04.891715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.324269Z digest=sha256:84b74b20d66cf582c81e85ab233602423d3640ce6563e30568af6ac5e3700df9

Observation 5953f810-81cf-4204-a698-f1788d9337c7 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.837972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.386780Z digest=sha256:310c3820fec41b9835bbc7808283565d2bd0df9654ce2ce14e8893bed507f3a3

Observation 89d98f07-0444-41d4-9ab2-b95c1ec61730 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.730723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.444609Z digest=sha256:7dae8dc288324b913fd3ec53c314213d6b80d142d964cbbe3f9468f19c330158

Observation c8fcdc46-aa79-4c0b-929c-710df3a816f8 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.570551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.490287Z digest=sha256:74f6215f009f6c58d3285f14d631d03923c11204d3c34ce5e0be556faac0a951

Observation 8074b593-43c7-4a83-a570-8071abf0a365 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.437871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.567217Z digest=sha256:889723ea3090a72594c7592a52804fd45b63a4ab4c142ecfa82f61483750d571

Observation 5825ffe0-a1f3-4f1c-af74-b36e3b47bd51 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:30:04.772792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:30:04.614811Z digest=sha256:cc17315b6115ba1637081b77f1903c2a811f299497465026a775193c4f98128c

Observation f75f7814-d361-42b0-a987-39bff45af3f1 · outbound

This paper cites MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation

Reference 1118

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:03.875329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:03.875329Z digest=sha256:5ec92d6bf54f5165f5da09379b2ce19e5570652ec3862499b650ee8cb24e16b0

Observation b3a50432-2799-4222-8072-24ab5db5925e · outbound

This paper cites The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:03.829875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:03.829875Z digest=sha256:5364a86b48af11ba4a93b37533924bb0d7f6817b685c32911f09d207253e7e70

Pith citing papers

Observation b555cb75-8a55-487b-8429-0ff88055c6cd · inbound

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas cites this paper.

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:32.360665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:32.360665Z digest=sha256:edcc1a4e938855770f04064267fcdd09f1a7efb4c1b2a060ae8d4af8b8fb7458

Observation 46c50fdc-078a-42df-9163-041d793d16ef · inbound

Robot builds a robot's brain: AI generated drone command and control station hosted in the sky cites this paper.

Robot builds a robot's brain: AI generated drone command and control station hosted in the sky ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:51:20.803345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:51:20.803345Z digest=sha256:418e868c653ef5dcbccf5fbf5f92d851bf28d6da1cecd5f2f4e5d5d9d374671b

Observation 3fdb24da-5aa2-4e20-9e2d-5d85a0516e6a · inbound

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models cites this paper.

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:21:22.599676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:21:22.599676Z digest=sha256:98b34500a1d56f2249e4b6b62ba0c6e5cd545cbf72e2f2b28bbdda0ad8501fdb

Observation 8845b873-a0d6-4b73-a5a1-001b7f67e920 · inbound

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents cites this paper.

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:15:28.008514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T18:15:13.924995Z digest=sha256:1ed72c5b62af1c07a92c61780550f81bad3ef4fe55a3f37c61dfe2b14bba1e7e

Observation 4ad5c576-5f75-422e-89b7-6473640b0e45 · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:06.935796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:1ee74eb999147afbb255b1dd4a36c01be47c430629ef5c89d478dfde089cba5d

Observation 7d50b9dd-91ab-491b-b104-8c950183688f · inbound

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification cites this paper.

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.271560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T17:32:06.864338Z digest=sha256:b654cf075beec6f16f95d59fd721a9a828bc55d0a2a9c1bfc5cfa2c46b4df7ca

Observation 9827d15f-5293-4b57-aa18-9223ba4c8115 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.778992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:987f812886d74944c198a634a6ad7c2ecf641dcd8a67f4a77ac8d596dbed648e

Observation eabd1c0e-4c1b-4689-a89f-5ed82c8c11d7 · inbound

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories cites this paper.

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:01:04.244734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T02:30:39.346257Z digest=sha256:1acb645b76847f0dfe1e081bcff1651a695647ae86b0c1db2f52bfdeb6904104

Observation eddb464f-bb32-48c8-91c2-7cbca15f67af · inbound

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench cites this paper.

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:52:43.035333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T17:49:01.198956Z digest=sha256:9ba1816c2eb8be41751bb2168bda702f11f7e373c2adcccf347445757c0adf38

Observation f78992f4-4ce3-462a-b6a9-68a132b93bb6 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.142766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:d1214f739fab78d9f6b60c1325fd8717bc117052f9795ce21d986cad930f90ef

Observation 78d3d367-84ca-450b-bab8-24e16520e1a2 · inbound

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents cites this paper.

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:57:53.363396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T23:54:15.987953Z digest=sha256:5d915da464d69f3d134d3f6a46b1ccf6e54f08e8ee3e19f69249ea08c216f8e3

Observation 87403219-8c32-45e1-bfd6-113e90b5ff56 · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.684486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:b4992f337db82ff0a17b3595e96dc7018164deb81aab052b9a4add510a0667e7

Observation dfaac76a-3c6a-4e6f-a532-dff57ebd65bf · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:35.697198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:35.697198Z digest=sha256:96a0d53de2f864bbd90235b930bd4fd99d851c53b68b09b5d939fe6401e98e8b

Observation 943dad94-76be-4e79-8e7e-ac03b93f6b2a · inbound

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence cites this paper.

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.672894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T12:13:58.111299Z digest=sha256:ed267873fb8f9b2439745ffbe75cf1e5b95ef2fbb53b985d44a28f762a006fff

Observation 15cc6769-5a69-4933-b731-70b6c788d4fb · inbound

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience cites this paper.

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-02T01:00:22.372434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:00:22.372434Z digest=sha256:2a8e8f7a76f40036817022eecb1292c4c0d84074e96e6cd1e5127d3c318d233b

Observation 88708047-8690-4df7-a06e-1ca8438e43fc · inbound

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis cites this paper.

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T14:26:52.414276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:26:52.414276Z digest=sha256:706cbedd3f5c2498f11c5905b6292c6e7db2259315ad8a0244833064ad583939

Observation 63a653a7-9bd5-43ab-b6a6-5c508d9d876e · inbound

Revibing Code from Papers: Reimplementing HCI Artifacts cites this paper.

Revibing Code from Papers: Reimplementing HCI Artifacts ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:46.242048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:24:46.242048Z digest=sha256:0f8e6236c5c92a624932b79ddc4b26debe92a901d2135c433eed88855b0581f6