Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 88 inbound Pith citation observations for arXiv:2504.01848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:44:21.776309Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
24
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2b502811-4942-428e-86a9-5d4dfe898868 · inbound
RExBench: Can coding agents autonomously implement AI research extensions? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1522369f-a2ac-412c-b1c6-3b185221915b · inbound
Kimi K2: Open Agentic Intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efe2d978-22e7-4464-9c6b-3e8158d9a37f · inbound
Evaluation and Benchmarking of LLM Agents: A Survey PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899640bd-8c91-443e-b784-8e78d7700c19 · inbound
How Far Are AI Scientists from Changing the World? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4fa4a3-6d43-4ab8-b363-b39db566f4f3 · inbound
TextQuests: How Good are LLMs at Text-Based Video Games? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a1eb17-837a-427d-8003-8dbac838c2ca · inbound
Reliable Weak-to-Strong Monitoring of LLM Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c8718a-f500-42f2-80f9-5492d1609b8e · inbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a6f9bc-c61a-4b15-9237-0f9624ff9a67 · inbound
Evalet: Evaluating Large Language Models through Functional Fragmentation PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36bd4472-ea6f-49ec-a2c2-f999fbe9401b · inbound
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c900501e-0855-41ed-9663-22471c67a678 · inbound
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fb2987-9169-495a-9dff-6ae253a63d3b · inbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a7171b-9d98-42ab-99bd-a8521a871200 · inbound
CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a526915-422d-4a4e-87ac-7c9c060f66d3 · inbound
ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b8e48d-112c-479c-aebb-8c185ebc7ed3 · inbound
AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6219e2b6-7f02-4af3-ab03-ff06febf5b53 · inbound
Kimi K2.5: Visual Agentic Intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65a82c89-567e-44b8-82cd-f489ab58f16c · inbound
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b70385-4c54-47c5-a026-9f61e93a434c · inbound
Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3807827d-d024-4bb4-b443-b9f7ecba8420 · inbound
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7b3e0b5-e862-4cb9-b439-5b8bf971e3d1 · inbound
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3244be22-c2e9-4041-bf4b-96e80ff2eef1 · inbound
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de791bc-76f8-471f-a1fc-dd9f7f0c117c · inbound
Effective Strategies for Asynchronous Software Engineering Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4649b9-597b-4055-9f3d-ff1525f6b71a · inbound
Towards Verifiable and Self-Correcting AI Physicists for Quantum Many-Body Simulations PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c9caef3-ccad-4c8e-8b46-4458463f9369 · inbound
FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84fbb1f2-2e1b-48dc-adfa-dd36c1863a5e · inbound
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbe5905c-3a12-4fb1-a183-b7cdfc493527 · inbound
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efd323f0-d183-439d-a4f0-42d244a41afb · inbound
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae98230-2d11-4743-a694-031fae0fe70e · inbound
In-Place Test-Time Training PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e372a679-17b7-4c3b-aae9-a8073b83ed22 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15e27c42-deb5-4175-92d3-446cf491b2b5 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab41ee1a-3c78-40c2-befb-1f1ae075a3d3 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33acf2d8-89e2-434a-8357-247ae7b0d8f8 · inbound
Evaluating LLM Agents on Automated Software Analysis Tasks PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac3f8fc7-f40b-4efe-8506-39a8a68c303a · inbound
Evaluating LLM Agents on Automated Software Analysis Tasks PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94eba6f1-b62f-41a1-810b-cec001c75c3a · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 986f8cfb-5031-4f01-9f0f-d1f5349f97c3 · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ec1fb92-a8c8-4ca0-8b67-545666e3c670 · inbound
Evaluation-driven Scaling for Scientific Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f61364c0-76bc-4c06-950e-5273f72fbc45 · inbound
Risk Reporting for Developers' Internal AI Model Use PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69210352-567a-4317-a497-3cdbf5d1be51 · inbound
ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e73d574-d455-423e-a21d-523752807a52 · inbound
ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0a159e8-6ab8-4324-9baf-ba3000436cc0 · inbound
AcademiClaw: When Students Set Challenges for AI Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 534dd0cc-5f37-46c8-bc61-7783b7d7672b · inbound
AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95fc485e-4eec-4430-91b9-8f5d212912bb · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5836dd27-66f0-4ea9-ae33-6867c4ed5422 · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e783be1-9655-437b-bdc2-25206bfc59d3 · inbound
ReproScore: Separating Readiness from Outcome in Research Software Reproducibility Assessment PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 379b324d-dd7f-40a3-abdd-e8396e169081 · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0d11633-060a-4552-86a7-6738dc104871 · inbound
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f52f1028-c800-40ec-a0c4-f19b8a7aee55 · inbound
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1277a15f-154f-4a4b-9feb-69607a1a3ee0 · inbound
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6cd443f-bfed-4275-b35f-2a5fccbb5d97 · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f874185-03d4-4125-88d8-5d71560a1adb · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5120d42e-0965-495e-ab34-3fc304e1390d · inbound
How Far Are We From True Auto-Research? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8100497b-ec48-480f-a262-17c4c2c432d0 · inbound
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48493cce-1ff8-4ce0-bb85-0e056ac80b1a · inbound
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bfff271f-f12a-44a9-95ef-1a83ee0055ab · inbound
Sibyl-AutoResearch: Autonomous Research Needs Self-Evolving Trial-and-Error Harnesses, Not Paper Generators PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d04ba36d-14ee-4c64-b588-5f1c1a9ae2ed · inbound
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0077aa23-1c0a-4026-82b1-6e3057abdd98 · inbound
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f467628d-f6b0-45de-8412-46d4fa7e8630 · inbound
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ce54e21-e98a-4489-8ca0-e38ccbf8d100 · inbound
From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0cdbdbb-f07d-48ae-b274-11c0b1372d11 · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation badb0b11-bb1d-4423-805a-39697d0a727b · inbound
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3547529-1383-4498-a1fe-f8ec1d4724c2 · inbound
TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b2b4bfd-f12b-4103-a973-152ea360d8b6 · inbound
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35e667e5-c9fa-4d36-bf7e-c7fbec662cfb · inbound
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af0b094a-2dee-4d8f-b78e-e3dd45b8b178 · inbound
Review the Code, Not the Story: A Vision and Protocol for Code-First Peer Review PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c723c8bd-596a-431d-90e8-73dceb3e2ec7 · inbound
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f97e3af-41bd-4c13-8fad-b0bce47408ff · inbound
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 137b7037-94fa-41e5-8a40-c4f541630a56 · inbound
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50701dc5-3c10-49ba-b09d-57c9fcd55005 · inbound
SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44b38372-6e5f-404c-8c00-73861e472074 · inbound
CEO-Bench: Can Agents Play the Long Game? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f89ac7ef-4e0c-45e3-8b49-535dd0b668e1 · inbound
CEO-Bench: Can Agents Play the Long Game? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1982cf77-4757-47e5-8aad-891d041ad855 · inbound
Skill Coverage: A Test Adequacy Metric for Agent Skills PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 957ae55a-8e66-40c3-85ed-a9856f6a2287 · inbound
PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2155120f-2cfd-4f8d-b2da-5636676073e9 · inbound
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d74561ea-7965-47dd-9250-0c525dc614c6 · inbound
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a44e1ac9-2656-460f-8ca3-30cb9f7f776a · inbound
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33ce8a9-c6b0-4f52-9f4c-e63d655b0fb4 · inbound
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67ee60a2-64b4-4719-915d-688e3e0a9d32 · inbound
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d32ecd3-33f2-41fd-8744-7bcb35fa6fd7 · inbound
Coding-agents can replicate scientific machine learning papers PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69ed3cf0-fca6-476f-abd8-13655ebe9598 · inbound
Coding-agents can replicate scientific machine learning papers PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60f2863-a29e-4ca3-b451-c72e77eb0e72 · inbound
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32c3176-36bd-40b9-b963-846344fa6c13 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3bc42954-0bac-450c-93e7-4b1ac1d1288b · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9eb065-8019-49f3-9a0e-8fe6752af78f · inbound
Towards Autonomous and Auditable Medical Imaging Model Development PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c201216e-df42-49a3-aa31-63698d94c0ab · inbound
Self-Improvements in Modern Agentic Systems: A Survey PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4180afd-16df-4f7d-8424-a0317e2d8806 · inbound
Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b21ee08-d16c-4454-b242-05b533ac0dc2 · inbound
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989a91c6-3816-4024-a079-b2f5b2a845fe · inbound
MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac05b38-ed7a-43cd-b0b9-8d02f0a39d79 · inbound
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3ae39e-671d-4674-a715-775352b9cce4 · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.