Pith. sign in

Paper Citation Record · LEDGER

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.28631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28631 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:54:35.911023Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54804b4f-9eb0-45f0-b60f-e6bfadbd4119 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review LitLLM: A Toolkit for Scientific Literature Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.811783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.811783Z digest=sha256:79130bce32d6180263e646f1424c2b5ebb0c5d96fa0a68e8f5138486d0114db8

Observation 3948c35b-512d-4fc0-802e-fad64cf328b4 · outbound

This paper cites Towards an AI co-scientist.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Towards an AI co-scientist

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.085990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.085990Z digest=sha256:2dc0d940f70f6f8c5db401355b42689eec376252d3e828eecabb710cc97c3e3f

Observation 7b4f0955-c022-4461-901f-79f3b7df63b2 · outbound

This paper cites Automated Algorithm Selection: Survey and Perspectives.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Automated Algorithm Selection: Survey and Perspectives

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.368864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.368864Z digest=sha256:1401963918c3d31735d6d38b84fb773bc367df88c09bd610bd4889374d77bf60

Observation 856ba543-d7a2-4d26-bf19-60ee6784b964 · outbound

This paper cites ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.900318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.900318Z digest=sha256:57add76b60358423c3602cc104c72b73e417b12a099ce135f55189a56423342e

Observation 0fb7e291-34af-482a-83d1-dc3f9e634416 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.012634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.012634Z digest=sha256:b40df5ec1391e70829d72add6f2d3531fc031d26dc507a41070bd81ee728d7b0

Observation e718b5ec-23ee-477c-b25e-534170c877eb · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.107501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.107501Z digest=sha256:3a54d6a6a2f030f845e7a544ce8204b1bacbdc91a9feb1de6eed28a2f821e4fa

Observation ef9c0efc-42a4-4cdb-952c-0dc5b32b13fc · outbound

This paper cites S., Bartley, N., Urbanowicz, R.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review S., Bartley, N., Urbanowicz, R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.227993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.227993Z digest=sha256:33eb0c7b14b717d673465506808213c5a5aad2502b6185f2ed1fd7beda13d782

Observation 38d1699a-f476-4fd9-8792-b7a0a14e5d22 · outbound

This paper cites Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.281409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.281409Z digest=sha256:fb444aedc4d2e689c5f6403e649d5ea4bbd7d68ece002df73705795f4713acbe

Observation a17570f3-d27a-4ef4-894a-7d79718cf68c · outbound

This paper cites Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.499762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.499762Z digest=sha256:c0b45da70a0bea0764077221a533779948d1b4b23ce4e8ce9e19295c93a6e36d

Observation 7a715175-85ea-4ed5-80a9-87f8a3baa9dc · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AI-Researcher: Autonomous Scientific Innovation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.646218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.646218Z digest=sha256:1fc48558225eecd361ad9251e19317a410a32de8c93fc723e8014c5fa0c68ba6

Observation 1ba4f510-578c-47c7-95db-22d733d556ac · outbound

This paper cites CycleResearcher: Improving Automated Research via Automated Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review CycleResearcher: Improving Automated Research via Automated Review

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.758261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.758261Z digest=sha256:ee654f540ff2728bd3ec248bee0409fa17ce9000b539bd2a3d3faf74e7fe73ce

Observation 030a3167-4768-451a-a90b-2dc9bb6a2219 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Neural Architecture Search with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.911023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.911023Z digest=sha256:483b2ae96265e554caf2fe4cffb11bbfed97207c2224dbe98c45c5c42a00d261

Observation 9b545b39-bfaf-41a9-af63-2d92395c6395 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Agent Laboratory: Using LLM Agents as Research Assistants

Reference 1976

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.395915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.395915Z digest=sha256:cf632e8581938d88c60355d8043be13fca8dfcf510ac1007129625f18b41239a

Observation c6bb8387-0b31-44ef-883d-8451320b7df3 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.833406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.833406Z digest=sha256:0482dfe98e5f4bbc7ea94bb371bcb0c37c330b093fc4f32f740ab3dd7da7731f

Observation 6ef9b525-834e-445f-8a9d-f8e3f98ee58f · outbound

This paper cites Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.631818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.631818Z digest=sha256:4616152160dc1b6422d6e36cc467870dababc19a52e991a29540d5aa6f84c839

Observation 1a9add53-6d65-48c7-a0e3-f7120dbbb5ab · outbound

This paper cites Crafting papers on machine learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Crafting papers on machine learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.512917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.512917Z digest=sha256:094d53d9096c5eed4070af57897016aeae5e39623d3c79f3c32add0fe1001656

Observation 9e2a4b00-8ce0-401b-aefc-c6667a5cf512 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review DARTS: Differentiable Architecture Search

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.776651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.776651Z digest=sha256:af753e4212b1f3739a47cb373c4ad14a5cf6124754b85c794afeb383935db7c7

Observation 76e81f91-4eb0-4b5d-b9a3-76b062e91b94 · outbound

This paper cites AgentReview: Exploring Peer Review Dynamics With LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AgentReview: Exploring Peer Review Dynamics With LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.236060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.236060Z digest=sha256:fcceaf6db4da15eb630bd58995db713035b980f4b9b830505ded1f866da7d7ae

Observation f9339de1-69b2-47e7-be8b-28c7717bcce9 · outbound

This paper cites K., Cucerzan, S., and Hwang, S.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review K., Cucerzan, S., and Hwang, S

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.952910Z digest=sha256:8410dee46b0a8487dae23d603361eb4ca0aabde758dc498d8119e965086c6473

Pith citing papers

No inbound Pith citation observations are available.