Pith. sign in

Paper Citation Record · LEDGER

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.28631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28631 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:54:35.911023Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54804b4f-9eb0-45f0-b60f-e6bfadbd4119 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review LitLLM: A Toolkit for Scientific Literature Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.811783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.811783Z digest=sha256:df4cf4e18bef717544e896b632ba9bda090cc4a0c7eda685b1090edbe1a5e4d8

Observation 3948c35b-512d-4fc0-802e-fad64cf328b4 · outbound

This paper cites Towards an AI co-scientist.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Towards an AI co-scientist

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.085990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.085990Z digest=sha256:cb22cbb59f93308d84b4303934081c42b87e67ca97dd0ec3e1e8b58db0af812c

Observation 7b4f0955-c022-4461-901f-79f3b7df63b2 · outbound

This paper cites Automated Algorithm Selection: Survey and Perspectives.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Automated Algorithm Selection: Survey and Perspectives

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.368864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.368864Z digest=sha256:ce88a39ef0b09eaef78ac0c9ef1ab370754dc44c53a284f113925b65a3e2364f

Observation 856ba543-d7a2-4d26-bf19-60ee6784b964 · outbound

This paper cites ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.900318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.900318Z digest=sha256:25da25a93752f1672047eb2dc45d3cfdc60e10eebed2cc99fa1c547ba98aaa3d

Observation 0fb7e291-34af-482a-83d1-dc3f9e634416 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.012634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.012634Z digest=sha256:56c07b56404bde3732d14b2d5bdde1ce11b3f7283f0ded0e18021daa9201187a

Observation e718b5ec-23ee-477c-b25e-534170c877eb · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.107501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.107501Z digest=sha256:f386c7808b637aa8d347265d0aa4ad0e7abf14c2427caea0cfb5e3f29d6dde28

Observation ef9c0efc-42a4-4cdb-952c-0dc5b32b13fc · outbound

This paper cites S., Bartley, N., Urbanowicz, R.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review S., Bartley, N., Urbanowicz, R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.227993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.227993Z digest=sha256:c77877aa141c947ae086f572d0da0454eb0287bef875e1786ea435f62d4c8e65

Observation 38d1699a-f476-4fd9-8792-b7a0a14e5d22 · outbound

This paper cites Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.281409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.281409Z digest=sha256:c8c36e81e81cb84b8743895f3716a3e28f25b1615e221de8ad156b3682c8faef

Observation a17570f3-d27a-4ef4-894a-7d79718cf68c · outbound

This paper cites Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.499762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.499762Z digest=sha256:c13bc84d0aa244e64f8f0bf6a35a2ee71742cfd0921c4d59e722c9b54a5d189b

Observation 7a715175-85ea-4ed5-80a9-87f8a3baa9dc · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AI-Researcher: Autonomous Scientific Innovation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.646218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.646218Z digest=sha256:c655ceda15ac99366ecb391bb5f1bfb55949fdbacfc56282d4d4f9314df50454

Observation 1ba4f510-578c-47c7-95db-22d733d556ac · outbound

This paper cites CycleResearcher: Improving Automated Research via Automated Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review CycleResearcher: Improving Automated Research via Automated Review

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.758261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.758261Z digest=sha256:91cc317a8f5a6b0e2fc2d42f151c8cadcd90b610bf482b0bbf2f8f0bf7664b6b

Observation 030a3167-4768-451a-a90b-2dc9bb6a2219 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Neural Architecture Search with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.911023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.911023Z digest=sha256:739e9e801b225e038898a9b67b5d610c875a9201ec019268897773093e5deebf

Observation 9b545b39-bfaf-41a9-af63-2d92395c6395 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Agent Laboratory: Using LLM Agents as Research Assistants

Reference 1976

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.395915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.395915Z digest=sha256:583d9b840d1d247642a41bdaf102afee850956f409d9695996beb8753af3e1a3

Observation c6bb8387-0b31-44ef-883d-8451320b7df3 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.833406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.833406Z digest=sha256:0e9c5f83b9d51da223001f4df9d5005e21b56aa12333462c55d76320c0a2808f

Observation 6ef9b525-834e-445f-8a9d-f8e3f98ee58f · outbound

This paper cites Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.631818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.631818Z digest=sha256:a8460ed0c1fec317d6619e83a0e3bed8b332ee7abe3c0cbf53d23a5f448c0ea3

Observation 1a9add53-6d65-48c7-a0e3-f7120dbbb5ab · outbound

This paper cites Crafting papers on machine learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Crafting papers on machine learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.512917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.512917Z digest=sha256:79d55c3e1c6c2a71759fa846a9838e3acdae6630607061d6ed58c159e6c5219c

Observation 9e2a4b00-8ce0-401b-aefc-c6667a5cf512 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review DARTS: Differentiable Architecture Search

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.776651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.776651Z digest=sha256:09eb83af10bd6b472be05085df70f97819d79d8dfd535659f73fbfe0c21abd37

Observation 76e81f91-4eb0-4b5d-b9a3-76b062e91b94 · outbound

This paper cites AgentReview: Exploring Peer Review Dynamics With LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AgentReview: Exploring Peer Review Dynamics With LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.236060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.236060Z digest=sha256:47a842d271eb701a6cd3c5e4b4056b88eeb113a8c8339cb7c483654019af556f

Observation f9339de1-69b2-47e7-be8b-28c7717bcce9 · outbound

This paper cites K., Cucerzan, S., and Hwang, S.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review K., Cucerzan, S., and Hwang, S

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.952910Z digest=sha256:f92832caba85ed1be78a017fb435609c805c9dc4450260f6e9f471ee7d258814

Pith citing papers

No inbound Pith citation observations are available.