Pith. sign in

Paper Citation Record · LEDGER

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2605.16616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.16616 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:59:40.519962Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.800152Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact13
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch19

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b166c4b-c536-4ebf-8064-ab6cd53f0cf6 · outbound

This paper cites GPT-4 Technical Report.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.113477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:411803842e7b3adcfd9120449ca066a19f5af392a6659432b022d802c5956402

Observation e31bd759-e308-415c-ad3c-a810caa9e9bf · outbound

This paper cites InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.403879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:25f5ae7a15c6c9617f9fc3ac4dfd8c0a6ad10619a201ea895cf65e5e9fe14477

Observation 82dd0b20-74ad-4494-859a-e198958f7e08 · outbound

This paper cites Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.396093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:e965c0137b4c5e7ce867aacfb9f4f75dc6ac28aef44ed58e6e13a108668f1ac1

Observation 58d4332f-1e00-4c80-b0c8-e5c3f61908f9 · outbound

This paper cites ChemCrow: Augmenting large-language models with chemistry tools.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ChemCrow: Augmenting large-language models with chemistry tools

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.099705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:ce14bcc7ba18e8d4446197ec684949b5cdf392ac314c6a864c889888074a7e5a

Observation 04fe23f3-d80f-4a3b-a8a1-d1102e13cd29 · outbound

This paper cites Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.407414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:31f3bd687797929cab72339f1950f5333e1bbebff935f0702de39b48c14eb599

Observation 92bd12e4-4819-4baa-8d3b-88a8a9a2219f · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.137430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:0b2c8ea0282c634aa44da11f2fe53db5cb25f56c24d141e83f5ee64816668f52

Observation 62a52266-d15c-461c-9fe8-8fde1897bea2 · outbound

This paper cites arXiv preprint arXiv:2505.19955(2025).

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility arXiv preprint arXiv:2505.19955(2025)

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.095253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:40e79bc01203e67310a48e94f89093c2972ea565fe4e149363bc291b02672177

Observation 9ae8b600-47c5-42ad-9d29-ecefac5a3aa4 · outbound

This paper cites ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.130932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:4f1b998592dfaf8b80e1ce2ac14a3b9cfe7b91172adbd1ae052db75de2788ece

Observation 59815ca7-b530-4067-b208-b88c3b69b009 · outbound

This paper cites The Value of Prediction in Identifying the Worst-Off.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility The Value of Prediction in Identifying the Worst-Off

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.104273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:970e60c6adb8f1ed7959afb4c3f65684f5baa9f40ef9fbf2d44f1a0f6d2da9c8

Observation ca2e8f01-3988-47cc-910c-fb5e7c0cc65e · outbound

This paper cites Josh Givens, Song Liu, and Henry WJ Reeve.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Josh Givens, Song Liu, and Henry WJ Reeve

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.400467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:f59a03b6064f47e1b7364d4e788efa460dec37b598de1cfed3e3b74db5adc08e

Observation 9925bb4d-f02a-4681-90b2-6c30b7d84333 · outbound

This paper cites Score Matching With Missing Data.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Score Matching With Missing Data

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.124752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:7c12a43452e2891bff3f7f1f781d5a783bc181a429e6e7fe12d9756343810554

Observation e027d510-67ac-46a0-b34e-d04081890451 · outbound

This paper cites Towards an AI co-scientist.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Towards an AI co-scientist

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.108786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:3fb37bfe11bbe91fbef7f30dcd5ee02f28443a15f5d777b5862eb21eb9664c0e

Observation f78992f4-4ce3-462a-b6a9-68a132b93bb6 · outbound

This paper cites ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.142766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:89508ce27b811c19e9141b2b6eb89684e8a69a87d9b652ee7e02d128c2226387

Observation b69a8c44-cb0f-4d33-b023-26623db79095 · outbound

This paper cites Idea2Plan: Exploring AI-Powered Research Planning.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Idea2Plan: Exploring AI-Powered Research Planning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:20:35.737828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:ca490e8a7a615ebbaefcd6dfe55167ecaa16f4630de11c87d6c78f5488b6f29b

Observation faa7c95a-d138-42c4-95fb-b654ee60a4e1 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.962049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:53a724d95b74f3fa25083d6702419dbef0219c2ff9fd45e34d8cb2931759fd94

Observation 1bd1277b-9871-45f5-a3f5-18f0af05c50e · outbound

This paper cites InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations).

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.427903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:6bd41f8b5d4d3c508e9dcd8deab80ffcbeb4e0561573c25e1bee26fbf3ac1778

Observation 9551d984-19d5-4aa2-a7df-748a34ed964b · outbound

This paper cites https://www.intology.ai/blog/ zochi-tech-report.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility https://www.intology.ai/blog/ zochi-tech-report

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.431562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:a9f89ab3679168cb963e4b5faa7e7865c27a5341fd624052f94adfee1f43057f

Observation e9e1b2ee-6db9-40a2-8b14-06cb12a01656 · outbound

This paper cites AIDE: AI-Driven Exploration in the Space of Code.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility AIDE: AI-Driven Exploration in the Space of Code

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:43.996922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:c4c1b62da27b81cd4ca5056118c38e9ca065cd1fd96c4e9584ed57e3a060c643

Observation 8eabc3ba-713e-42fd-99c0-7591fd6dfcd7 · outbound

This paper cites Machine Learning115, 5 (2026).

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Machine Learning115, 5 (2026)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.410566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:832e3f2401d9acde1097444a09c32c2a0e2a9935bd92b1e423b2bab1518022ce

Observation c9f67b27-7dd0-4b06-a420-6ccc2042afe6 · outbound

This paper cites Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.025076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:9455f5c9fba1c7b1ae822e8e211f8e57a1713f157a2f5f40e620f43a8f00899c

Observation 9b55ba47-63d8-44a0-baa9-3762a2d948e9 · outbound

This paper cites Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.031083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:1c2dd87ba2ab196aeecd0068607ba46cfda986a42344a12d6f4d449ae7e34591

Observation 0b27f761-0e83-4968-b7d1-d84806b97925 · outbound

This paper cites Mlr-copilot: Autonomous machine learning research based on large language models agents.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Mlr-copilot: Autonomous machine learning research based on large language models agents

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.991684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:04fd523b823fba286ff560b9f87447e54be2fdabec5de98b5858c4adf059cdcc

Observation 8ece0086-ce39-4559-a478-f253f869ddb4 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.036698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:91931117def5ba285225a0b1de17485e3d6a652dac7341f8d87d6f758398bc21

Observation 300413d4-f31d-4a8e-9275-467a4e57aec8 · outbound

This paper cites Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.013482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:c9ffb2bb8f7b08d4632e301a07e3e5d0c6f874abdde0b5c890a698ae573f64a7

Observation bb6bf2b1-c166-4a23-b57c-4b04936e025b · outbound

This paper cites MLGym: A New Framework and Benchmark for Advancing AI Research Agents.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.085167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:936a9f62ac20309b0884c89c1a32f5a72e960232195944fc77a45ee8688e2970

Observation 71594fa6-0b39-467d-a124-ec2601e58981 · outbound

This paper cites AI Idea Bench 2025: AI Research Idea Generation Benchmark.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility AI Idea Bench 2025: AI Research Idea Generation Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:43.974081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:e4cf3f497cd380bcb66e155c0764c546d016dac1ea4b06d061be5b42f965271c

Observation 0756e5a5-7897-4a1e-b28d-37b3d15f0a15 · outbound

This paper cites Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:43.967487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:28893f5fd144ad3bacb5d3695efb81c508c2e66ab44d7b7401479c8b1d41051a

Observation b3a48133-c903-433c-a2f9-319a0e231ab7 · outbound

This paper cites Minju Seo, Jinheon Baek, Seongyun Lee, and Sung Ju Hwang.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Minju Seo, Jinheon Baek, Seongyun Lee, and Sung Ju Hwang

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.423970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:aeb0fe97f8a65a8cba8b201f7084219614970626cc179ba79bc3d97cdd20b1ce

Observation 5a6e87bc-99c4-4d28-829b-d371a70de1b4 · outbound

This paper cites Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.074843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:114f7f98035b6c7c5072f5033d87cbf4a0884287dd626f11888454daf25e51ad

Observation 0e380be1-f4b2-4628-a75c-6e3b81f58de2 · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-24T01:14:19.028324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:a0b8ce8207c0275ce0a2ae40e9c9db1f2720dcc89d78a2c08b57812a93adafcf

Observation deb8ea53-0f27-4d6b-8000-42b9fc854511 · outbound

This paper cites Conformal Prediction as Bayesian Quadrature.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Conformal Prediction as Bayesian Quadrature

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.007931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:61c049ddb7bf578bbb1314cf19976f57bfcb9522259ee3d462ed14be9839c2d3

Observation f52f1028-c800-40ec-a0c4-f19b8a7aee55 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.002213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:55ae51dcadb749d5c179492154d918e86d6b06cd57df5e01f7abfaa5bc234874

Observation 3e720c4e-45a8-4000-ac09-a942e9132d31 · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility AI-Researcher: Autonomous Scientific Innovation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.047519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:1e872ad2b6286cf651c2735a536cd29881d0667237afda961c0a36685c2b2f71

Observation d3fa7a11-82fd-420e-a2f6-a7b46370ccf0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:e8926c71321ea6c523af3dd5899774af0ba070da5230ff9119ef08c2a502674f

Observation 2cc68a4f-c023-4249-8f2b-22f56d670f2d · outbound

This paper cites Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.413939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:a81ce3659feb240baad61bbf19558042286b379e5718db77170d70d5169c81c5

Observation cedccc7d-ee4a-4bdb-97c2-a53372c98bab · outbound

This paper cites CycleResearcher: Improving Automated Research via Automated Review.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility CycleResearcher: Improving Automated Research via Automated Review

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.069785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:7f8fd8e5b283fe6909ee9bbc8410a95dd257bb07247809548ba0a9766d7dcd3e

Observation a28acc7c-07de-4b87-8203-a5a354d290f8 · outbound

This paper cites CollabLLM: From Passive Responders to Active Collaborators.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility CollabLLM: From Passive Responders to Active Collaborators

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.080101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:ce63ae909ffd8398ba582fec42eabc678e2e5a2a105d0e7a821da7bc3c0b8d16

Observation f49d13c1-7280-445b-b733-8e2cccab9811 · outbound

This paper cites SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.090179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:780ae811f32682ea531eeb571bbec561ad1f61b9b694b82f6ebc41d3449e03c1

Observation 7ec5ac8a-d6b0-46f5-b3da-3512722da8d8 · outbound

This paper cites SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.052307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:cdf3bec387469fa80eb23ae9a2f7f9f8a79c39129ad81a5d90b92ec4e8a8ed6e

Observation 9aff99f0-4850-4eb6-b1f0-dde0fc3ad75f · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.057404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:271df38ca15feb8f51fb9e8b8b8d70c28b270628f447a0ad34bebcc6d4b670e6

Observation f88f855f-8dd1-40e0-8d56-2bd040ecf91c · outbound

This paper cites AI Scientists Fail Without Strong Implementation Capability.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility AI Scientists Fail Without Strong Implementation Capability

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.063892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:9e0994824788b2cb3caafdddcc7d9ae314814e2877215665dcbc03369607887d

Observation 5f89d451-e4d1-40b6-bac6-7eecae9d90ae · outbound

This paper cites experiment_name.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility experiment_name

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.417577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:ea6d164610b71ea53231a50c68eb4e3dd0cc212fae68a28b1535bce06205b435

Observation 41d5decd-2981-446f-815a-19628d7fe986 · outbound

This paper cites target".

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility target"

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:03:59.420844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:9916a80fc5c06ccf342599a0c066ca2878c9562bd2bc4e8e5f62b18355138d33

Pith citing papers

Observation 05e10731-bada-4fa4-8566-3aca3ab0fcc2 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

Reference 190

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.801393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:d289a4e3728ab6dff888783fc9ba236d2ca141f4a76fabd6f8cb1fc51174f251