REVIEW 4 major objections 5 minor 19 references
This paper claims that a three-LLM review protocol can reliably rank autonomous AI Scientist systems, and that on a 75-paper benchmark, FARS papers score more than twice as high as four competing frameworks—while Gemini and Claude, but not
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:51 UTC pith:LUFT4BSF
load-bearing objection First head-to-head benchmark of four AI Scientist systems with three LLM reviewers, but the FARS 2x claim is not supported: the reference standard is FARS's own papers on FARS's own proposals, the non-FARS scores sit at the floor, and the validation loop is circular. the 4 major comments →
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that when four autonomous AI Scientist frameworks (Sakana AI v1, Sakana AI v2, CycleResearcher, Data-to-Paper) are run on the same 15 research proposals, the FARS system's papers are substantially higher quality, and that this ranking is reliably recoverable by automated LLM review. FARS itself generated complete papers for each of the 15 proposals; the four frameworks produced 60 additional papers, for 75 total. Three independent LLM reviewers scored every paper on four 1–5 dimensions—Originality, Scientific Rigor, Clarity, and Significance—and a synthesis score was computed. FARS led on every dimension and reviewer, with more than double the score of the best c
What carries the argument
The load-bearing mechanism is the multi-model LLM review protocol: three independent frontier LLMs (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) each score every paper on a 1–5 scale across four dimensions—Originality, Scientific Rigor, Clarity, and Significance—which are combined into a synthesis score. The benchmark's controlled design is the second pillar: 15 FARS proposals, converted to each framework's input format via custom wrappers (with Data-to-Paper receiving datasets extracted from GitHub repositories), so all systems work from the same research questions; Sakana v1 and v2 outputs are merged using an LLM-based reconciliation system. Spearman rank correlation among reviewers serves as
Load-bearing premise
The load-bearing premise is that the 15 FARS-generated papers are a neutral, high-quality reference and that 1–5 scores assigned by LLMs measure genuine research quality—neither premise is calibrated against human expert judgment.
What would settle it
Have human experts in AI safety and machine learning blind-score the same 75 papers on the same rubric, then check whether FARS still leads by roughly 2× and whether Gemini–Claude agreement matches human inter-rater agreement; alternatively, rerun the benchmark with 15 proposals authored by a different system or by humans and see if FARS's advantage disappears.
If this is right
- If FARS's 2× margin is real, its multi-agent architecture—specialized ideation, planning, experiment, and writing agents sharing a workspace—appears to be a substantially better template for autonomous research than linear or loop-based pipelines.
- Multi-model LLM review of this kind could be deployed as a cheap, fast first-pass filter for AI-generated papers, reserving human review for borderline or high-stakes cases.
- The strong Gemini–Claude consensus suggests that agreement between two independent LLMs can substitute for human ground truth when ranking papers within this task distribution.
- GPT-5.4's divergent rankings imply that single-reviewer automated evaluation is unreliable; benchmarks and quality-control pipelines should use at least two agreeing models.
- The measured cost and throughput trade-offs (Data-to-Paper at $9 and 20 minutes per paper, CycleResearcher at 10 minutes per paper) give practitioners concrete data for choosing a system when quality is not the only constraint.
Where Pith is reading between the lines
- Because FARS authored both the proposals and the reference papers, the 2× margin may partly reflect a home-turf advantage; rerunning the benchmark with proposals authored by a different system or by humans would test whether the gap persists.
- Agreement between Gemini and Claude does not by itself establish validity—both models could share correlated biases—so calibrating the scores against human expert judgments is necessary before treating automated review as a true replacement for peer review.
- A natural extension is to check whether LLM review scores predict downstream outcomes such as citations, replication success, or acceptance by human reviewers; if they do, the protocol becomes a genuine measurement instrument rather than a comparative gauge.
- Another testable extension is to run the same 75 papers through additional LLM families (including open-weight models) to see whether the Gemini–Claude cluster is unique or a general property of frontier models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a benchmarking protocol for autonomous AI Scientist systems, running four frameworks (Sakana AI v1/v2, CycleResearcher, Data-to-Paper) on 15 research proposals and comparing their outputs against 15 FARS-generated reference papers. Three LLM reviewers (GPT-5.4, Gemini, Claude) score each paper on 1–5 scales, and a synthesis score is computed. The central findings are: (i) FARS papers score substantially higher than all competing frameworks, with more than 2× higher means on Gemini and Claude; (ii) Gemini and Claude show strong agreement (ρ = 0.907), which is presented as validation of the automated evaluation; (iii) cost and runtime differ substantially across systems. The paper concludes that this establishes the first quantitative benchmark for AI Scientist systems and demonstrates that multi-model LLM evaluation is a scalable, consistent framework.
Significance. The paper undertakes a substantial empirical effort: it executes four real AI Scientist frameworks over 15 controlled proposals, produces 75 papers, and reports cost/efficiency trade-offs. If the evaluation method were externally calibrated and the reference set were neutral, this would be a useful benchmark for a rapidly growing area. However, the central comparison is undermined by three intertwined problems: FARS supplies both the proposals and the reference papers (home-turf effect); the reviewer 'validation' is circular because the synthesis score is an aggregate of the same reviewers; and the non-FARS scores sit almost entirely at the floor of the 1–5 scale, making the reported 2× advantage an artifact of measuring against the minimum. No human review or independent quality labels are used, so the scores have no external calibration. The paper's claimed contribution as a benchmark is currently unsupported by its protocol; what remains is a descriptive case study with uncontrolled confounds.
major comments (4)
- [§3 and §4, Finding 1] The benchmark is structurally circular: the 15 proposals are FARS-generated, and the 15 reference papers are FARS outputs ("These 15 FARS-generated papers serve as benchmark references"). Non-FARS systems are therefore evaluated on another system's home turf, with Data-to-Paper even receiving datasets from FARS-associated repositories. The claim that "FARS benchmark papers significantly outperform" competing systems (Finding 1) is partly a consequence of this design. A fair comparison would require proposals drawn from a neutral source and reference papers not produced by the system being ranked; otherwise the 2× gap cannot be attributed to quality rather than task familiarity.
- [§3.2 and §4, Finding 2] The validation of the multi-model reviewers is circular. The “synthesis score” is described as a consensus aggregate of the three LLM reviewers, and Finding 2 reports strong correlations of Gemini and Claude with that synthesis score (ρ = 0.961) as evidence of reliability. By construction, an aggregate correlates highly with its components; the correlation does not indicate that the reviewers measure research quality. The weak GPT-5.4 agreement (ρ ≈ 0.32) shows that the reviewers disagree substantially, and high Gemini–Claude agreement could arise from shared model priors or shared prompt design. No human review or independent external quality labels are used, so the claimed reliability has no external calibration.
- [Table 2] The headline “more than 2×” result is undermined by a severe floor effect. In Table 2, nearly every non-FARS paper scores 1/5 on Gemini and Claude, and many score 1 on all reviewer models; FARS means are 2.14–2.47. A ratio computed against a near-minimum baseline (e.g., 2.47/1.07) does not quantify a quality gap—it only shows that non-FARS outputs are below the measurement floor. The table provides only overall scores, not per-dimension distributions or per-proposal spreads, so it is impossible to assess whether the floor effect is driven by a single egregious issue or by uniformly poor outputs. The claim of a “substantial quality gap” is therefore not supported by the reported data.
- [§3.2] The synthesis score is never precisely defined. The protocol describes two different evaluation systems—one with dimensions Structural Coherence, Citation Analysis, Novelty Assessment, and Methodological Rigor, and another with Originality, Scientific Rigor, Clarity, and Significance—but Table 2 and Finding 2 refer to a single “Synthesis” score without specifying which dimensions are used, how the three LLM outputs are aggregated, or whether a separate synthesis prompt/model is applied. This omission prevents replication and makes the validation claim unfalsifiable.
minor comments (5)
- [§3] Typo: “Fo our” should be “For our”; also the reference list contains informal author names like “Technion-Kishony-lab” that should be formatted consistently.
- [§2.4] The sentence “CycleResearcher achieves an average quality score of 5.36 on a standardized 5-point scale” is internally inconsistent; a 5-point scale cannot produce a mean above 5. Please correct the scale or the number.
- [§3.1] The LLM used for merging Sakana v1/v2 outputs is not specified (model, prompt, criteria). Since merged papers are what get evaluated, this step should be described in sufficient detail for replication.
- [Table 2] The mean row shows “20/1/1/1” for CycleResearcher, which appears to be a formatting error for “2.0/1/1/1”. Please fix and double-check all numeric entries.
- [General] The paper does not include a data or code availability statement. Given that the dataset of 75 papers and evaluation prompts are central to the claims, releasing them would materially improve reproducibility and allow community checks of the floor-effect and home-turf concerns.
Circularity Check
LLM-reviewer validation is self-referential and the FARS benchmark is home-turf; the central quality claims are uncalibrated.
specific steps
-
self definitional
[§4 Finding 2; §3.2 Evaluation Protocol; Figure 2 caption]
"Synthesis scores (aggregate across GPT-5.4, Gemini, and Claude reviewers). ... Both Gemini (ρ = 0.961, p < 0.001) and Claude (ρ = 0.961, p < 0.001) correlate extremely strongly with the synthesis score, demonstrating that the synthesis metric successfully captures the Gemini-Claude consensus."
The synthesis score is defined by the paper as an aggregate of the GPT-5.4, Gemini, and Claude reviewer scores (Figure 2 caption). A reviewer's correlation with the synthesis is therefore a self-correlation with a quantity containing that reviewer's own scores; high ρ is largely guaranteed by construction and cannot establish that the scores measure 'genuine quality signals.' Gemini–Claude agreement (ρ=0.907) is agreement between two same-kind LLM instruments, not calibration against human peer review or any external quality label. The paper nonetheless uses this internal consistency as evidence of 'reliability of automated evaluation,' closing the validation loop on itself.
-
other
[§3 Experimental Setup and Protocol; §4 Finding 1; Table 2]
"For our benchmarking study, we use a curated set of 15 research proposals generated by the FARS (Fully Automated Research System) system. ... These 15 FARS-generated papers serve as benchmark references for comparison, representing a quality standard produced by a mature automated research system."
The system that is ranked first (FARS) authored both the test proposals and the 'benchmark reference' papers against which all other systems are compared. Ranking four non-FARS frameworks on FARS-authored proposals and against FARS-authored references cannot separate a genuine quality gap from a home-turf effect, e.g., proposals and phrasings that are naturally matched to FARS's own pipeline. Moreover, Table 2 shows that non-FARS Gemini/Claude/Synthesis scores are almost all exactly 1, so the headline 'more than 2× higher' (FARS 2.47 vs. 1.0) is computed against the minimum scale point and is a floor artifact rather than a calibrated measurement. No human review, released papers, or external quality label anchors the ranking.
full rationale
The paper's central claims are empirical rather than formally derived, and no self-citation chain or imported uniqueness theorem is used. However, two load-bearing inferences reduce to the paper's own inputs. First, the 'validation' of the multi-LLM review instrument is circular: each reviewer's very high correlation with the synthesis score is a self-correlation, since the synthesis score is the aggregate of those same reviewers, and the 'reliability' claim rests on inter-LLM agreement rather than on any external calibration (human scores, known quality labels, or machine-checked ground truth). Second, the benchmark reference is FARS itself: FARS generated the 15 proposals and FARS-generated papers are the standards of comparison, so the 'FARS significantly outperforms' finding is a home-turf measurement. Table 2's floor effect—all non-FARS systems scored 1 on Gemini/Claude/Synthesis for nearly every proposal—makes the 2× ratio uninterpretable. These two issues jointly undermine the abstract's claims of a 'rigorous benchmarking protocol' and 'reliable automated evaluation,' but they are partial circularity rather than a fully definitional collapse, hence score 7.
Axiom & Free-Parameter Ledger
free parameters (3)
- synthesis score aggregation rule =
unspecified (uses GPT-5.4, Gemini, Claude scores)
- LLM evaluation rubric/prompt template =
not released
- ideas-per-proposal for Sakana (k=3) =
3
axioms (5)
- domain assumption LLM 1–5 scores are valid proxies for scientific paper quality.
- ad hoc to paper FARS papers are a legitimate quality reference standard.
- domain assumption Inter-rater agreement implies reliability.
- domain assumption The 15 proposals are a neutral task set.
- ad hoc to paper LLM-merged Sakana outputs fairly represent Sakana v1/v2.
invented entities (1)
-
Synthesis score
no independent evidence
read the original abstract
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all competing frameworks, achieving mean scores of 2.14--2.47 on a 1--5 scale compared to 1.00--1.87 for other systems. Notably, FARS scores are more than 2$\times$ higher than the next-best systems on Gemini and Claude evaluations. We find strong agreement among Gemini and Claude ($\rho$ = 0.907, $p < 0.001$), and both correlate extremely strongly with the synthesis score ($\rho$ = 0.961, $p < 0.001$), validating the reliability of automated evaluation. However, GPT-5.4 exhibits weaker agreement ($\rho \approx 0.32$), suggesting it evaluates papers using different criteria. These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality.
Figures
Reference graph
Works this paper leans on
-
[1]
Agarwal, S., Sahu, G., Puri, A., Laradji, I. H., Dvijotham, K. D. J., Stanley, J., Charlin, L., and Pal, C. LitLLM: A Toolkit for Scientific Literature Review. arXiv preprint 8 Submission and Formatting Instructions for ICML 2026 arXiv:2402.01788,
Pith/arXiv arXiv 2026
-
[3]
Towards an ai co-scientist.arXiv preprint arXiv:2502.18864,
Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., Palepu, A., Sirkovic, P., Myaskovsky, A., Weissenberger, F., Rong, K., Tanno, R., et al. Towards an ai co-scientist.arXiv preprint arXiv:2502.18864,
-
[5]
H., Neumann, F., and Trautmann, H
Kerschke, P., Hoos, H. H., Neumann, F., and Trautmann, H. Automated Algorithm Selection: Survey and Perspectives. arXiv preprint arXiv:1811.11597,
-
[9]
Liu, R. and Shah, N. B. ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Re- viewing. arXiv preprint arXiv:2306.00622,
-
[10]
T., Foerster, J., Clune, J., and Ha, D
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.arXiv preprint arXiv:2408.06292,
-
[11]
Novikov, A. et al. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery. arXiv preprint arXiv:2506.13131,
-
[12]
S., Bartley, N., Urbanowicz, R
Olson, R. S., Bartley, N., Urbanowicz, R. J., and Moore, J. H. Evaluation of A Tree-Based Pipeline Optimization Tool for Automating Data Science. InProceedings of the Genetic and Evolutionary Computation Conference 2016,
2016
-
[13]
arXiv preprint arXiv:2509.01659,
-
[15]
Starace, G. et al. PaperBench: Evaluating AI’s Ability to Replicate AI Research. arXiv preprint arXiv:2502.16069,
-
[16]
AI-Researcher: Autonomous Scientific Innovation
Tang, J., Xia, L., Li, Z., and Huang, C. AI-Researcher: Autonomous Scientific Innovation. arXiv preprint arXiv:2505.18705,
-
[17]
Weng, Y ., Zhu, M., Bao, G., Zhang, H., Wang, J., Zhang, Y ., and Yang, L. CycleResearcher: Improving Auto- mated Research via Automated Review.arXiv preprint arXiv:2411.00816,
-
[19]
Zoph, B. and Le, Q. V . Neural Architecture Search With Re- inforcement Learning. arXiv preprint arXiv:1611.01578,
-
[1976]
Agent Laboratory: Us- ing LLM Agents as Research Assistants
9 Submission and Formatting Instructions for ICML 2026 Schmidgall, S., Su, Y ., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Liu, Z., and Barsoum, E. Agent Laboratory: Us- ing LLM Agents as Research Assistants. arXiv preprint arXiv:2501.04227,
Pith/arXiv arXiv 2026
-
[1997]
T., Lu, C., Hu, S., Lu, C., Foerster, J., Clune, J., and Ha, D
Yamada, Y ., Lange, R. T., Lu, C., Hu, S., Lu, C., Foerster, J., Clune, J., and Ha, D. The AI Scientist-v2: Workshop- Level Automated Scientific Discovery via Agentic Tree Search.arXiv preprint arXiv:2504.08066,
-
[2000]
Li, L., Xu, W., Guo, J., Zhao, R., Li, X., Yuan, Y ., Zhang, B., Jiang, Y ., Xin, Y ., and Dang, R
Morgan Kaufmann. Li, L., Xu, W., Guo, J., Zhao, R., Li, X., Yuan, Y ., Zhang, B., Jiang, Y ., Xin, Y ., and Dang, R. Chain of Ideas: Revolutionizing Research via Novel Idea Development with LLM Agents. arXiv preprint arXiv:2410.13185,
-
[2018]
Crafting papers on machine learning
Langley, P. Crafting papers on machine learning. In Langley, P. (ed.),Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stan- ford, CA,
2000
-
[2023]
DARTS: Differentiable Architecture Search
Liu, H., Simonyan, K., and Yang, Y . DARTS: Differentiable Architecture Search. arXiv preprint arXiv:1806.09055,
-
[2024]
AgentReview: Exploring Peer Review Dynamics With LLM Agents
Jin, Y ., Zhao, Q., Wang, Y ., Chen, H., Zhu, K., Xiao, Y ., and Wang, J. AgentReview: Exploring Peer Review Dynamics With LLM Agents. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,
2024
-
[2025]
K., Cucerzan, S., and Hwang, S
Baek, J., Jauhar, S. K., Cucerzan, S., and Hwang, S. J. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. InPro- ceedings of the 2025 Conference of the Americas Chapter of the Association for Computational Linguistics,
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.