Pith. sign in

Paper Citation Record · LEDGER

PaperBench: Evaluating AI's Ability to Replicate AI Research

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 88 inbound Pith citation observations for arXiv:2504.01848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.01848 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 88 of 88 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.776309Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2b502811-4942-428e-86a9-5d4dfe898868 · inbound

RExBench: Can coding agents autonomously implement AI research extensions? cites this paper.

RExBench: Can coding agents autonomously implement AI research extensions? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:37:08.857659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T07:33:39.675929Z digest=sha256:7ffa0fcd0fdce65b8ed4af23bb8ba688d5ba83fa4fb3f5bcbddfa09339bf241c

Observation 1522369f-a2ac-412c-b1c6-3b185221915b · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:ecc5ab1d5df01eef7434398a825777d982119491fc50ab1ff232a0fce5f10435

Observation efe2d978-22e7-4464-9c6b-3e8158d9a37f · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.776309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.776309Z digest=sha256:7b008fcecb7bba15eb33f189a38ed216fbf2c30d5853be43aa68614ce55fae77

Observation 899640bd-8c91-443e-b784-8e78d7700c19 · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:15.021219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:15.021219Z digest=sha256:b6684bf7f16f146fffcfdcfded8d87b3d277f48a19a84a4bf5e8b14b9944d26a

Observation cf4fa4a3-6d43-4ab8-b363-b39db566f4f3 · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.675992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.675992Z digest=sha256:4e817d5c1cb46a9daacbbc7919ad710100b9dd0928d0b918947e4e53a6965dc3

Observation 40a1eb17-837a-427d-8003-8dbac838c2ca · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:54.201712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:54.201712Z digest=sha256:deff9d5627d87b5132b617ed05af312261fcee42b7d8db63e0d77b8b40300e33

Observation b1c8718a-f500-42f2-80f9-5492d1609b8e · inbound

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries cites this paper.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.654747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.654747Z digest=sha256:18bc043623d6c3afa07c29040a383a15b25735ea524a658b1331ad67d9662d04

Observation 22a6f9bc-c61a-4b15-9237-0f9624ff9a67 · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:01:39.890501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:b04cff414bcf9bb78449f08a8246a7e03e434e22bdfed1e0cab9742ec8ae765c

Observation 36bd4472-ea6f-49ec-a2c2-f999fbe9401b · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:06:32.139371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:01ad167f6c03b396744a34966f653547dd7c2b516a5d53702aca8b748b071d54

Observation c900501e-0855-41ed-9663-22471c67a678 · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:24.531675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:24.531675Z digest=sha256:f56477301f71e5fe241a4fbdfffb5ff6b66d2bc80e2aa52d5e2b069f1f12057f

Observation 96fb2987-9169-495a-9dff-6ae253a63d3b · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:44.150906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:44.150906Z digest=sha256:2b25fafc7418430f2782be61978b49d422be8d339291a5f9a874905c93d3c6a1

Observation 70a7171b-9d98-42ab-99bd-a8521a871200 · inbound

CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases cites this paper.

CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:10:48.468533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T03:10:21.188635Z digest=sha256:dace258fc8b565cdcd4f848fd14503e81c22db7368a9b944e42c671b2d9a7855

Observation 4a526915-422d-4a4e-87ac-7c9c060f66d3 · inbound

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction cites this paper.

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:34:51.330668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:34:51.330668Z digest=sha256:8ae5facc1ae341b05b67db3d840044e75c7f5e2fcf1ff5460000bd049fad5160

Observation 95b8e48d-112c-479c-aebb-8c185ebc7ed3 · inbound

AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows cites this paper.

AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:52:53.260889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:52:48.443287Z digest=sha256:a4d0219f4c48142c74645bf9513f3e7da64324b4db9a035e32da90a40e75fca5

Observation 6219e2b6-7f02-4af3-ab03-ff06febf5b53 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:524e8524adf03f8ca37a2f9801f3102843fe316fd518c9d5075c8bc5b6cdd0de

Observation 65a82c89-567e-44b8-82cd-f489ab58f16c · inbound

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights cites this paper.

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T05:17:18.391571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:17:18.391571Z digest=sha256:4207ba03c3975861e105fbcc61baa2bf001c4de36d3e556e07c1485719b5aba6

Observation 25b70385-4c54-47c5-a026-9f61e93a434c · inbound

Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches cites this paper.

Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:57:24.419477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:55:48.203213Z digest=sha256:6f7182f13c12546bb8a845ab0b35923c00e946da455a8936e00bda0ede7acaab

Observation 3807827d-d024-4bb4-b443-b9f7ecba8420 · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:02:06.911840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:b7d12d8bf56af8d5aefd0dd7ccd713007bb56730dacbc6380b1bd3cae1a3ac63

Observation d7b3e0b5-e862-4cb9-b439-5b8bf971e3d1 · inbound

Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"? cites this paper.

Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:52:53.443887Z digest=sha256:4d6b43d0f0c85d00992d6b3e50a97acee46f22dc1aebe823bf360c4f76cd10f5

Observation 3244be22-c2e9-4041-bf4b-96e80ff2eef1 · inbound

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking cites this paper.

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:09:20.123350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:09:20.123350Z digest=sha256:7454ff8210a35ac9f186b8ece270b56562fcd2e9a22d90656b865b6186df360d

Observation 3de791bc-76f8-471f-a1fc-dd9f7f0c117c · inbound

Effective Strategies for Asynchronous Software Engineering Agents cites this paper.

Effective Strategies for Asynchronous Software Engineering Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T20:49:09.477849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:49:09.477849Z digest=sha256:33f0a3c5e38d92bf10cd4cd1e3fcf8e6eefff0f010ca7d73edff976f52940a16

Observation ae4649b9-597b-4055-9f3d-ff1525f6b71a · inbound

Towards Verifiable and Self-Correcting AI Physicists for Quantum Many-Body Simulations cites this paper.

Towards Verifiable and Self-Correcting AI Physicists for Quantum Many-Body Simulations PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:33:21.677835Z digest=sha256:92d54e79a587037c9b7602f590e6b691c59ebe77e4c9f2e4b087ed7a5102742c

Observation 8c9caef3-ccad-4c8e-8b46-4458463f9369 · inbound

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification cites this paper.

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T17:32:06.864338Z digest=sha256:c698fd622128bd825169aa74c65e115686181e28bf4c86562f0f03143404aa2f

Observation 84fbb1f2-2e1b-48dc-adfa-dd36c1863a5e · inbound

RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers cites this paper.

RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:26:34.666529Z digest=sha256:4255be947fa5539c5be39b8b1c460d851882df2d69ea05cfcfd0fb885b2b9af2

Observation cbe5905c-3a12-4fb1-a183-b7cdfc493527 · inbound

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery cites this paper.

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:43:51.525392Z digest=sha256:f40e17df0296a587cf3d0482ffee095b31857492e271d6ec6e1c3419ca86f710

Observation efd323f0-d183-439d-a4f0-42d244a41afb · inbound

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery cites this paper.

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T09:23:40.328143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:23:40.328143Z digest=sha256:3d16020c3ceb6fef3da329e0514b5780cf2cf26c2a96cc0ca7190cb94a4b04dc

Observation eae98230-2d11-4743-a694-031fae0fe70e · inbound

In-Place Test-Time Training cites this paper.

In-Place Test-Time Training PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:07:47.174513Z digest=sha256:d700c0f8e098a15c866a09444c658fe24c2e024ca62b563d8282e6c50bf516a9

Observation e372a679-17b7-4c3b-aae9-a8073b83ed22 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:32285f1d1af9bf539debfd0a2994e3c9300ff22fec7c775a573ce64eeea5a8d3

Observation 15e27c42-deb5-4175-92d3-446cf491b2b5 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.451453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.451453Z digest=sha256:afd7a5c1bc5e0e8c85f7ff4ca5fa6258b0142c3671058138ff68e63831ec88cf

Observation ab41ee1a-3c78-40c2-befb-1f1ae075a3d3 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.369991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.369991Z digest=sha256:ef8fe45213dc021bd9ab162de0db1d35c1d3a4097eb48a85a67c6e54d94b9262

Observation 33acf2d8-89e2-434a-8357-247ae7b0d8f8 · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:38:12.144252Z digest=sha256:9310f2d71c0b9658dc35940b968159e84b62ece3f76c5160552403d9b747ff8b

Observation ac3f8fc7-f40b-4efe-8506-39a8a68c303a · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:44.960426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:44.960426Z digest=sha256:cca1d744e0f9434df187228c34b57ac00f6dae69148aade2bab921e89ffb8996

Observation 94eba6f1-b62f-41a1-810b-cec001c75c3a · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:2449034642a2c84ee1a8c7fa1b1cc4efe232e09ca439915c896bef6028603f6d

Observation 986f8cfb-5031-4f01-9f0f-d1f5349f97c3 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.819303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:439bd79bdc102ad3d70df4ad32b0ef6a0bdadd94f9ad07babeafd93916fd2244

Observation 4ec1fb92-a8c8-4ca0-8b67-545666e3c670 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:5c798cb4ff6913aef49726cff20a98d598fba0a4e79f8a3b0d19cf7b2ddae295

Observation f61364c0-76bc-4c06-950e-5273f72fbc45 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:3011da12963108a0c3e2716356fadc0bd2ed63868d8a7401ade6a8f4d83dc203

Observation 69210352-567a-4317-a497-3cdbf5d1be51 · inbound

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review cites this paper.

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T01:46:17.836724Z digest=sha256:7e111cfa60b5c174214dcd977c1fa920daa698c62532a6f31314b83ed7625c97

Observation 9e73d574-d455-423e-a21d-523752807a52 · inbound

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review cites this paper.

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T17:17:42.608991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:14:18.558837Z digest=sha256:407a3d62744e84ac97313fb8bec923698bc4f1ff97a1dfeadb18300ec2a883a3

Observation f0a159e8-6ab8-4324-9baf-ba3000436cc0 · inbound

AcademiClaw: When Students Set Challenges for AI Agents cites this paper.

AcademiClaw: When Students Set Challenges for AI Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T19:24:29.696454Z digest=sha256:48b11634dde2fc5967ce3d9e31d26d0319a6f8a9485aef5e92a221c304dafc1d

Observation 534dd0cc-5f37-46c8-bc61-7783b7d7672b · inbound

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents cites this paper.

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T22:02:47.160731Z digest=sha256:0c29ada5fdb5f01b8547529c0030c04a54870c78ac6e5be32af690f14db42e11

Observation 95fc485e-4eec-4430-91b9-8f5d212912bb · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:23:34.514071Z digest=sha256:ec09d9472447bbb87a70ed18437ee6ba756cb96064c62c3fb492e6566f2bfea4

Observation 5836dd27-66f0-4ea9-ae33-6867c4ed5422 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T04:51:17.519200Z digest=sha256:d6da5210626a58872b27c77d89e3a2232f08490457812147e0e4a378693b744f

Observation 8e783be1-9655-437b-bdc2-25206bfc59d3 · inbound

ReproScore: Separating Readiness from Outcome in Research Software Reproducibility Assessment cites this paper.

ReproScore: Separating Readiness from Outcome in Research Software Reproducibility Assessment PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:27:14.564913Z digest=sha256:e3667ba593ddf6ca8df9e5c6b7f3b4ad5a5c29ba26fb7d90bfeabc3dc2f49a28

Observation 379b324d-dd7f-40a3-abdd-e8396e169081 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:09:35.570850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:07b17432ac3589e5b91782f6abe8b28928b1c99d843159a1a1b513c982ee7518

Observation c0d11633-060a-4552-86a7-6738dc104871 · inbound

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design cites this paper.

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T18:58:53.828922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T18:58:11.197587Z digest=sha256:52ef449ecdd9aaca1177e37897252e207981196bcf77ddbf6f5542f94f71d9b6

Observation f52f1028-c800-40ec-a0c4-f19b8a7aee55 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.002213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:0e5300959fe9eed91bf5989939ff480f3d589e3790d3ff09035a5a176853b31e

Observation 1277a15f-154f-4a4b-9feb-69607a1a3ee0 · inbound

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery cites this paper.

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T20:32:45.507229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:29:27.398862Z digest=sha256:cf21c749dad6b080592e927d66ba069b66127b6c349c899b1c17412236306e24

Observation f6cd443f-bfed-4275-b35f-2a5fccbb5d97 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:28:17.244214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:f1d94f2d4ff1cf2c63a9ee2e724a61a3ce8c7ebeeb07e85f5936c329224d6c4a

Observation 0f874185-03d4-4125-88d8-5d71560a1adb · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:23.262252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:f6e0e808e8a0dba4f7180e213bb1c0171193afaf1aa34dcd48f5dd4b90714d8d

Observation 5120d42e-0965-495e-ab34-3fc304e1390d · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T09:58:11.246139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:572439d279eedd6c786393a958d6c7ef31e07ca014925e7c492b0a0053e654b9

Observation 8100497b-ec48-480f-a262-17c4c2c432d0 · inbound

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work cites this paper.

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T04:13:56.508484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:10:37.172329Z digest=sha256:2c4d405a711db90b3df9865ec2244363cd40c1c1becdfa7ff03feca6820d730b

Observation 48493cce-1ff8-4ce0-bb85-0e056ac80b1a · inbound

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work cites this paper.

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T09:44:45.904629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:42:30.305592Z digest=sha256:5e6ce987de1da73e8ff2c14866362ad262dc971c96050dbd29f66d6c84cb6768

Observation bfff271f-f12a-44a9-95ef-1a83ee0055ab · inbound

Sibyl-AutoResearch: Autonomous Research Needs Self-Evolving Trial-and-Error Harnesses, Not Paper Generators cites this paper.

Sibyl-AutoResearch: Autonomous Research Needs Self-Evolving Trial-and-Error Harnesses, Not Paper Generators PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T02:05:55.413691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T02:05:05.934682Z digest=sha256:3f9b0af6571875dd5710a9dc9db205d3567b7b8f0c4120282ae87a31ce9c5e5d

Observation d04ba36d-14ee-4c64-b588-5f1c1a9ae2ed · inbound

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery cites this paper.

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 109

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:50:21.769927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T04:46:43.679185Z digest=sha256:42b621810d7b5a418748253b4d7f6acc10b40be82a7e92dc392bea43590e6918

Observation 0077aa23-1c0a-4026-82b1-6e3057abdd98 · inbound

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence cites this paper.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:23:59.092437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T21:19:03.281629Z digest=sha256:7c65de05e6e04f721a474a6c7e3f7a6317d46e821975705a0189ce516df6098e

Observation f467628d-f6b0-45de-8412-46d4fa7e8630 · inbound

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems cites this paper.

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:33:32.787883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T15:32:21.737028Z digest=sha256:bdfcb52ea67e14ee7378c3a2a15bcfe33ab5f95ca9a10ebd316bb17047761d7b

Observation 4ce54e21-e98a-4489-8ca0-e38ccbf8d100 · inbound

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence cites this paper.

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:23:24.635069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T12:13:58.111299Z digest=sha256:8827ebd9154f5a664bc5198837b114fff9d629327edcc9a5db6afca9fbf9a3ab

Observation d0cdbdbb-f07d-48ae-b274-11c0b1372d11 · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.152849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:bcc60e531abe42385ae1e003d6b3182a457d9c09056850ad6e9cd8de85ae4599

Observation badb0b11-bb1d-4423-805a-39697d0a727b · inbound

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment cites this paper.

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:16:49.260505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T05:35:12.027697Z digest=sha256:2a3b6ea7e4a75db24227d2326cc4f25c611da25e69739cac6477a9566f7f8653

Observation b3547529-1383-4498-a1fe-f8ec1d4724c2 · inbound

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework cites this paper.

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:56.495517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:55:59.002171Z digest=sha256:35ccad4b50833e5c39dc4a2447e41f2cee56fbd70905a9ac57bb08e48fb0afc4

Observation 0b2b4bfd-f12b-4103-a973-152ea360d8b6 · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:33:15.810705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T08:24:42.412763Z digest=sha256:514b0f9d6f672926fe90dcdb4e0e86e014475eca4ec47f6283fc6769a672f8c9

Observation 35e667e5-c9fa-4d36-bf7e-c7fbec662cfb · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:39:16.541721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-04T00:30:00.449270Z digest=sha256:b2d0d74cea5d0b76ca6078f2e56c4ba550b9ebcbe6eca7667cece8242f28ab55

Observation af0b094a-2dee-4d8f-b78e-e3dd45b8b178 · inbound

Review the Code, Not the Story: A Vision and Protocol for Code-First Peer Review cites this paper.

Review the Code, Not the Story: A Vision and Protocol for Code-First Peer Review PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T18:57:17.031265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:45:03.399754Z digest=sha256:607d9804626ba59712dca5b0682dabb7e80d8fa5c90c528f555a942e4e483df2

Observation c723c8bd-596a-431d-90e8-73dceb3e2ec7 · inbound

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline cites this paper.

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T22:01:20.844772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:76783af3bc86cf815d0ca74a76c96e1240c68cfe3514c64505cb6fd69b7a7d5d

Observation 5f97e3af-41bd-4c13-8fad-b0bce47408ff · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:07:37.246876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:431a50c8b6929b3573d79da6e9bf4f4f9b137ee541ed69b9d79f839da27844be

Observation 137b7037-94fa-41e5-8a40-c4f541630a56 · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 152

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T09:40:46.983091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:457a10f5ab746b0ebee9ffce77146604f74580b2488d9e36ce63ad1cac8bb35b

Observation 50701dc5-3c10-49ba-b09d-57c9fcd55005 · inbound

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents cites this paper.

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:09:00.158258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T23:40:53.439549Z digest=sha256:da4f7e3db6f3e51d9638648042e0298ab0ad64d4b207b6abb1723f061363e82b

Observation 44b38372-6e5f-404c-8c00-73861e472074 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.689019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:0173f1ae08a27f7b00e39ad2a20f718649b93e656f6a02345cd2e034c76c4b14

Observation f89ac7ef-4e0c-45e3-8b49-535dd0b668e1 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:19.550132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:19.550132Z digest=sha256:53724d7abdbcf6d0f6dbdbb59df78192291b890d618b29dc7801975466ec0afe

Observation 1982cf77-4757-47e5-8aad-891d041ad855 · inbound

Skill Coverage: A Test Adequacy Metric for Agent Skills cites this paper.

Skill Coverage: A Test Adequacy Metric for Agent Skills PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T13:40:57.378439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:d5a266aa7373644518baf3939385fa39ab4ada7c6d79a144ca7681444e15a810

Observation 957ae55a-8e66-40c3-85ed-a9856f6a2287 · inbound

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation cites this paper.

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:20:00.267468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:47:47.892277Z digest=sha256:5b5c19c1121b35a399cb2a2caf47adfbc815b0768eea88705fac8c0dc042513a

Observation 2155120f-2cfd-4f8d-b2da-5636676073e9 · inbound

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy cites this paper.

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T17:50:00.427072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:24:55.361203Z digest=sha256:3a120e14d2516a1ea84b85ee3e4c70117e5fcbdd2bf972e50718e80641e2ec75

Observation d74561ea-7965-47dd-9250-0c525dc614c6 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:57.846088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T00:13:14.940915Z digest=sha256:f4c71ec0864574a449fe5d4d279a0f6d5cd50d0c1bc9584f4f889366010249f3

Observation a44e1ac9-2656-460f-8ca3-30cb9f7f776a · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T12:32:39.706055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:32:39.706055Z digest=sha256:db101cbd4789e5d94bcd58e7f8ed5277495ec3dcd5bab6ea1b615ef739de9bed

Observation b33ce8a9-c6b0-4f52-9f4c-e63d655b0fb4 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:29:56.680099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:1513cb79f99646ca1b545f4ce76a6d4bdbc4b0d2b7dabcda9640daa11b164b62

Observation 67ee60a2-64b4-4719-915d-688e3e0a9d32 · inbound

Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration cites this paper.

Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:24:22.543275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:15:29.120280Z digest=sha256:0ba3f685dc79063c5953609682829c84f078c6150da123105ed161ed81b2f6bf

Observation 4d32ecd3-33f2-41fd-8744-7bcb35fa6fd7 · inbound

Coding-agents can replicate scientific machine learning papers cites this paper.

Coding-agents can replicate scientific machine learning papers PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T13:08:07.696988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-03T13:07:26.748722Z digest=sha256:0ba932cf58690a909f703967a6150fe42d929d9fcb634475282ce6201d496fae

Observation 69ed3cf0-fca6-476f-abd8-13655ebe9598 · inbound

Coding-agents can replicate scientific machine learning papers cites this paper.

Coding-agents can replicate scientific machine learning papers PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T07:13:13.354959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T07:13:13.354959Z digest=sha256:eed79c4885b202ba204ff5d4b0d8b800a489e68371a53bab8678d1c80192a12c

Observation c60f2863-a29e-4ca3-b451-c72e77eb0e72 · inbound

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research cites this paper.

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T06:02:20.361005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:02:20.361005Z digest=sha256:bff388a326a7ccb9a6b24fdb41a073109d1ddb8f47e189b8f4c6fb0f60bd3195

Observation e32c3176-36bd-40b9-b963-846344fa6c13 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.196203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:03f8a38180cc6d61ba0f8983bff91e225e68ba040e02f4057a2ee7d4ee7989c7

Observation 3bc42954-0bac-450c-93e7-4b1ac1d1288b · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:48.856448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:48.856448Z digest=sha256:a4d47bb44a16de5886501a0c748eab1f9c500b918b27d9cfe9eeb2d95e54c577

Observation ca9eb065-8019-49f3-9a0e-8fe6752af78f · inbound

Towards Autonomous and Auditable Medical Imaging Model Development cites this paper.

Towards Autonomous and Auditable Medical Imaging Model Development PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T11:05:51.130839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:05:51.130839Z digest=sha256:ad8b53503d2cec56a76c65aa1a75a8c9fd57ac1e698292c05dcf4d4959a1dd8b

Observation c201216e-df42-49a3-aa31-63698d94c0ab · inbound

Self-Improvements in Modern Agentic Systems: A Survey cites this paper.

Self-Improvements in Modern Agentic Systems: A Survey PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-02T06:35:59.308317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:35:59.308317Z digest=sha256:719a005645a1d066d8d9d34a91e0710804e138e1583e7e62da526e9af4ec5aee

Observation e4180afd-16df-4f7d-8424-a0317e2d8806 · inbound

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience cites this paper.

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:00:22.885499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:00:22.885499Z digest=sha256:e9057ebe804e87c0f136ae37bf16ee629eb499c34be0f507ce01f03757a44049

Observation 8b21ee08-d16c-4454-b242-05b533ac0dc2 · inbound

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering cites this paper.

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T01:39:48.664815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:39:48.664815Z digest=sha256:0f075264a02b603a31c91ff7e19683ea664a995975513644665f2d15904fcaf3

Observation 989a91c6-3816-4024-a079-b2f5b2a845fe · inbound

MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment cites this paper.

MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:12.166627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:31:12.166627Z digest=sha256:ec23d023e74add3286f47c2b38d03de09574f2de76face3ec828a0afc8d98030

Observation 2ac05b38-ed7a-43cd-b0b9-8d02f0a39d79 · inbound

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities cites this paper.

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T16:49:31.209853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:49:31.209853Z digest=sha256:bfb8fbedb66bae714e23c9c7306c7d9b5fb3f3b48b3c2a0fbea4e102f9798e4e

Observation 6c3ae39e-671d-4674-a715-775352b9cce4 · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:43.695729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:43.695729Z digest=sha256:7632a57df89062e52dbe3c426e5aa372469be8f71527997f7db8155c1e1e1b6a