{"id":"1015fa41-22b8-4d87-8a1d-f59baa50756d","arxiv_id":"2507.23276","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.","lead":"AI Scientist systems built on large language models can already draft research papers, but this survey argues they are still far from making discoveries that change the world, and it organizes the field into four capability levels. A generalist reader will find a structured map of what automated scientists can and cannot do today, plus evidence that current systems still fall short of publishable-quality research.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative 'not good enough' claim rests on an uncalibrated AI reviewer with no human baseline; without calibration or a defined percentile reference, Tables 4–5 cannot bear the argument's weight.","rationale":"The reader's weakest assumption correctly identifies the DeepReviewer-14B calibration problem as the key vulnerability in the paper's new empirical evidence. My reading of the full manuscript confirms that §5.3 is the only direct evaluation of complete AI-generated manuscripts, and its methodology section does not report any human validation, calibration, or percentile-baseline definition. The concern is therefore load-bearing for the strong claim that current systems 'cannot independently produce scientific artifacts that meet established standards.' At the same time, the concern does not justify rejection. The survey's qualitative synthesis, the capability-level framework, and external benchmark evidence (Table 2) independently support a more general conclusion that current AI Scientist systems face major implementation and verification gaps. The paper also transparently flags that its 28-paper sample 'may be curated,' which is an honest limitation. Thus the appropriate response is to keep the reader's CONDITIONAL verdict: the paper should release the evaluation artifacts or temper the specific numeric and percentile claims, but the survey's overall message remains credible. I would not move the verdict because the reader already conditioned on exactly this issue, and the central qualitative thesis is supported by multiple independent strands of evidence.","tokens_in":37466,"tokens_out":3304,"duration_ms":38456,"concrete_test":"Take the 28 assessed papers (or a random subset of at least 10) and have 3–5 independent human experts with relevant domain expertise rate them on the same soundness, presentation, and contribution scales and the same defect checklist. Compute agreement (e.g., Spearman correlation and mean-score differences) between DeepReviewer-14B and the human ratings, and compare DeepReviewer's percentile outputs against a human-rated baseline of accepted workshop papers. If human ratings also place all five systems below a 5/10 threshold and reproduce the defect ordering, the concern is resolved; if human ratings diverge substantially (e.g., accepted papers score ≥6/10), Tables 4–5 need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §5.3—that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards'—is carried almost entirely by Table 4 and Table 5, both generated by DeepReviewer-14B (Zhu et al., 2025). This reviewer was built by authors overlapping with the present paper's author list, and the paper gives no calibration against human expert ratings, no inter-rater reliability, and no description of the reference distribution used to compute the 'Percentile' column. If DeepReviewer-14B is systematically harsh, the sub-4.63/10 averages and the 100% 'Experimental Weakness' rate could largely reflect reviewer bias or template artifacts rather than genuine defects in the assessed papers. The paper itself notes that the 28 papers are publicly available and 'may be curated,' so the sample may not be representative of typical output; with only 2–10 papers per system, the percentile estimates are also noisy. External benchmarks in Table 2 do support implementation-related gaps, but they do not directly establish the full-manuscript-quality claim made in §5.3. Because this section is the only place the paper directly evaluates complete AI-generated manuscripts, the headline conclusion depends on an unvalidated instrument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a four-level capability framework for AI Scientist systems (knowledge acquisition, idea generation, verification and falsification, and evolution) and reviews representative methods and benchmarks for each level. It also presents a new empirical evaluation in Section 5.3, where 28 publicly available papers from five AI Scientist systems are scored by DeepReviewer-14B, yielding average ratings below 4.63/10 and a 100% incidence of 'Experimental Weakness' across the papers. On this basis, the paper concludes that current AI Scientist systems cannot independently produce scientific artifacts that meet established standards for high-quality scientific communication. The survey closes with limitations of foundation models, research-capability gaps, ethical considerations, and future directions.","tokens_in":37828,"tokens_out":2954,"duration_ms":33137,"significance":"The survey is comprehensive and timely, and the proposed capability framework provides a useful organizing structure for a rapidly evolving field. The paper compiles external benchmarks (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) that credibly demonstrate implementation and verification gaps in current systems. The new evaluation of full AI-generated manuscripts, if properly validated, would be an important contribution. However, the quantitative 'not good enough' claim currently rests on an uncalibrated AI reviewer with overlapping authorship, and the paper does not provide the statistical context needed to interpret the scores. With appropriate calibration and error analysis, the central conclusion could be made solid; as it stands, the evidence is not yet load-bearing in its reported form.","major_comments":[{"comment":"The central claim that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards' rests entirely on ratings from DeepReviewer-14B (Zhu et al., 2025), yet the paper provides no calibration against human expert judgments, no inter-rater reliability statistics, no error bars, and no description of the reference distribution that defines the 'Percentile' column. Because DeepReviewer shares two authors with this manuscript, systematic bias cannot be ruled out. This is load-bearing, since Tables 4–5 are the only direct evaluation of complete AI-generated manuscripts in the paper.","section":"§5.3, Tables 4–5"},{"comment":"The evaluation sample consists of only 28 papers, with 2–10 per system, and the paper itself notes in the Table 4 caption that publicly available papers 'may be curated and therefore may not fully represent the typical output of each system.' The paper does not report confidence intervals or significance tests, so the ranking across systems and the aggregate 'not good enough' conclusion are not robust to sampling variability. The authors should either enlarge the sample, report uncertainty, or soften the claims to match the evidentiary strength.","section":"§5.3, sample and statistics"},{"comment":"The external benchmarks in Table 2 (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) directly support claims about weak implementation and verification capabilities, but they do not measure the holistic manuscript-quality dimensions (soundness, presentation, contribution) that Tables 4–5 assess. The paper should explicitly separate these two kinds of evidence and state that the manuscript-quality conclusion depends on the unvalidated DeepReviewer evaluation, not on the benchmarks.","section":"§4.2 vs. §5.3"}],"minor_comments":[{"comment":"The model name 'CinicalBERT' should be 'ClinicalBERT' (Huang et al., 2019).","section":"§2.1"},{"comment":"Figure 3 contains garbled text including long '/uni00000029/uni00000044/...' sequences and an unreadable table header ('Type Total citations Avg. citations'); this likely reflects a rendering or encoding error that must be fixed.","section":"Figure 3"},{"comment":"The system name 'The AI Scientist' is typeset with non-standard spacing (e.g., 'A I Sc i e n t i s t'), making some sentences difficult to read; please use consistent formatting.","section":"Throughout"},{"comment":"In the sentence 'current AI Scientist systems not only struggle with scientific execution but also stuck with clearly articulating their research findings,' the phrase 'but also stuck' should be 'but also struggle.'","section":"§5.3"},{"comment":"The 'Percentile' column is undefined; the caption should state the reference distribution (e.g., percentile relative to what population of papers or scores).","section":"Table 4"},{"comment":"The paper does not describe how DeepReviewer-14B computes the overall 'Rating' from the three sub-scores (Soundness, Presentation, Contribution), nor whether the 0–10 scale permits non-integer intermediate values; please clarify the scoring protocol.","section":"§5.3 / Table 4"},{"comment":"Several references have inconsistent formatting, including 'KABENAMUALU et al., 2023' in all caps and 'preprent' in Skarlinski et al. (2024); please proofread the bibliography.","section":"References"},{"comment":"The phrase 'prospect-driven review' is used but never defined; consider adding one sentence explaining what makes the review 'prospect-driven.'","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"I note that the authors of this manuscript overlap with the authorship of DeepReview (Zhu et al., 2025), the model used to produce the paper's central quantitative result. While overlap itself is not disqualifying, I would ask the editor to ensure that any revised version discloses this relationship clearly and, ideally, validates DeepReviewer-14B against human expert ratings or an independent AI reviewer. There is also a question of scope: the manuscript is currently structured as an arXiv-style survey, and the journal should consider whether the empirical Section 5.3 meets its standards for methodological rigor after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey does something genuinely useful: it gives the AI Scientist field a shared vocabulary (four capability levels: knowledge acquisition, idea generation, verification/falsification, evolution) and a sober status check grounded in third-party benchmarks. The literature coverage is broad and fair, and the Figure 3 citation analysis—papers with implementation details get more citations—is a nice, small empirical point. The GitHub repo for tracking the field is a plus. The authors also deserve credit for stating the obvious but often dodged conclusion: current systems are far from autonomous discovery, and the bottleneck is verification and falsification, not idea generation.\n\nThe genuinely new piece is Section 5.3: the evaluation of 28 papers from five AI Scientist systems using DeepReviewer-14B, with defect-category percentages and percentile rankings. That is original and worth reporting. But the stress-test note lands. DeepReviewer-14B is a model built by overlapping authors, there is no calibration against human expert ratings, no inter-rater reliability, and no description of the reference distribution behind the 'Percentile' column. The sample is small, likely curated, and the code/data aren't released. Table 5's 100% 'Experimental Weakness' rate could partially reflect reviewer bias or template artifacts. The paper itself notes the curation caveat, which is honest but doesn't fix the problem.\n\nHere's the thing: the central conclusion does not actually depend on that evaluation. The external benchmarks in Table 2 (MLE-Bench 16.9%, PaperBench 26%, SciReplicate-Bench 39%) independently show that verification and implementation are weak. So the paper's main message survives. What should change is the presentation of Tables 4 and 5: they should be labeled as a pilot study with an unvalidated instrument, or the authors should release the artifacts and do a human calibration study. As is, the specific numeric claims (e.g., 'cannot independently produce scientific artifacts that meet established standards') are over-strong when inferred only from DeepReviewer scores.\n\nAll that said, this is a serious, coherent survey that will be useful to researchers and funders tracking the area. It deserves a real referee process, not a desk reject, with the empirical section flagged for major revision.\n\nRecommendation: engage with it. Send to peer review.","headline":"A useful, well-organized status survey of AI Scientist systems whose central 'not good enough' claim is right, but whose own new measurement leans on an uncalibrated AI reviewer and should be treated as illustrative, not definitive.","tokens_in":38255,"tokens_out":1168,"would_cite":true,"duration_ms":15571,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current AI Scientist systems cannot yet produce papers that pass scientific muster; the strongest system averages 4.63/10 under an automated reviewer, and experimental weakness appears in every assessed paper.","keywords":["AI Scientist","large language models","automated scientific discovery","hypothesis generation","verification and falsification","AI peer review","scientific capability framework"],"falsifier":"Have a panel of human experts blind-review the same 28 papers, along with human-written accepted papers, and compare their scores and defect categories with DeepReviewer-14B's; if humans rate the AI papers near parity with accepted human work, or if DeepReviewer-14B gives equally low scores to strong human papers, the central claim is unsupported.","tokens_in":37247,"feed_emoji":"🔬","tokens_out":4513,"duration_ms":45706,"temperature":0.7,"pith_summary":"This survey argues that today's large-language-model-based AI Scientist systems are far from ready to reshape scientific research. The authors propose a four-level capability ladder—knowledge acquisition, idea generation, verification and falsification, and evolution—and use it to organize the field. Their new empirical evidence is a quality assessment of 28 publicly available papers produced by five leading systems, scored by an AI reviewer model. The best average rating is 4.63 out of 10, and every assessed paper shows experimental weakness. The paper concludes that current systems cannot independently produce artifacts meeting high-quality scientific communication standards, and that the bottleneck is verification and implementation, not idea generation alone.","feed_headline":"AI-generated papers top out at 4.63/10 in new quality check","feed_subtitle":"Survey finds experimental weakness in all 28 assessed papers; five AI research systems fall short of scientific standards.","key_machinery":"The organizing device is a four-level capability framework that defines what a mature AI Scientist must do: acquire knowledge from literature, generate feasible novel hypotheses, verify and falsify them through experiments, and evolve from feedback. The load-bearing empirical instrument is DeepReviewer-14B, an AI reviewer model that rates papers on soundness, presentation, and contribution and also lists defect categories; Table 4 and Table 5 use it to score 28 papers from five systems. The framework matters because it converts the vague question \"how far are we?\" into a checkable milestone list, and the reviewer scores supply the quantitative answer.","core_discovery":"The central claim is that no current AI Scientist system can autonomously carry out the full research loop well enough to produce work that would pass genuine scientific scrutiny. The evidence is twofold: on implementation benchmarks, even the strongest language models score low at reproducing or executing research code, and on the paper's own evaluation, all five surveyed systems average below 4.63 out of 10, with \"Experimental Weakness\" flagged in 100% of the 28 papers. The authors attribute the gap to limits of the foundation models—hallucination, costly knowledge updating, and catastrophic forgetting—and to underdeveloped research abilities in feasibility assessment, rigorous experimentation, and long-term planning. Their proposed remedy is an explicit \"evolution\" capability, where systems improve through self-reflection, external feedback, and structured collaboration, before they can be expected to produce ground-breaking discoveries.","pith_inferences":["The paper's scores come from an AI reviewer, which is itself a language model; a natural next test is to check whether the same reviewer would give equally low percentiles to human-written papers, which would reveal whether the scale is harsh overall rather than specific to AI output.","One testable extension is to run the evaluation longitudinally: if iterative review-feedback cycles raise later-generation papers above the current 4.63/10 ceiling, that would support the authors' claim that evolution, not raw model scale, is the missing ingredient.","The framework suggests a practical benchmarking protocol—score any new AI Scientist on all four levels separately—so progress claims can be compared across systems instead of relying on anecdotal acceptance at workshops.","Because the 28 papers are publicly available and therefore likely curated toward higher quality, the true typical output of these systems may be even weaker than the already low average scores reported."],"forward_implications":["If the assessment is right, workshop acceptance of AI-generated papers is weak evidence of scientific maturity, since the same systems' full outputs rate far below typical accepted work.","Verification and implementation, not idea generation, are the binding constraints; improving code execution and experimental design should come before claims of autonomous discovery.","The four-level ladder implies that an AI system must demonstrate all levels—including evolution through feedback—before it is called a scientist, giving the field a concrete evaluation target.","Deploying current systems without safeguards would flood peer review with low-quality artifacts, so the paper's proposed detection, labeling, and human oversight mechanisms become prerequisites.","The review's category data, with 100% of papers showing experimental weakness and 96.4% showing methodological unclarity, provides a checklist that future systems can be measured against to track real progress."],"supporting_citations":[{"why":"Supplies DeepReviewer-14B, the AI reviewer model that produces the quality scores and defect categories in Tables 4 and 5.","marker":"Zhu et al., 2025"},{"why":"One of the five AI Scientist systems evaluated, and the canonical example of an LLM agent claiming fully automated open-ended discovery.","marker":"Lu et al., 2024"},{"why":"Another evaluated system, CycleResearcher, which also supplies the iterative review-feedback mechanism the paper highlights as an evolutionary path.","marker":"Weng et al., 2025"},{"why":"An evaluated system that reportedly passed workshop-level peer review, providing the strongest apparent counterexample the paper tests and rejects.","marker":"Yamada et al., 2025"},{"why":"An evaluated system, Zochi, which the paper credits with tree-search planning but which still ranks below the quality threshold.","marker":"Intology, 2025"},{"why":"An evaluated system, HKUSD AI Researcher, providing another data point for the low-score distribution in Table 4.","marker":"Jiabin et al., 2025"},{"why":"MLE-Bench shows low end-to-end implementation accuracy, supporting the paper's claim that verification is a major bottleneck.","marker":"Chan et al., 2024"},{"why":"PaperBench shows low replication scores, evidence that current models cannot turn conceptual understanding into working experimental code.","marker":"Starace et al., 2025"},{"why":"SciReplicate-Bench demonstrates the best models reach only 39% execution accuracy, quantifying the implementation gap.","marker":"Xiang et al., 2025"}],"fun_headline_variants":["AI scientists score below 4.63/10 on all 28 papers in new survey","Survey: all 28 AI research papers fail experimental quality check","Five AI scientist systems all underperform in quality review","AI-generated science not yet credible: all papers show weak experiments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that DeepReviewer-14B's quality scores are valid and unbiased measures of scientific merit, since the paper does not calibrate the model against human expert judgments or test whether its harshness affects all systems equally.","fun_headline_variants_meta":{"raw":{"variants":["AI scientists score below 4.63/10 on all 28 papers in new survey","Survey: all 28 AI research papers fail experimental quality check","Five AI scientist systems all underperform in quality review","AI-generated science not yet credible: all papers show weak experiments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000369,"raw_usage":{"total_tokens":1952,"prompt_tokens":889,"completion_tokens":1063,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":988}},"tokens_in":505,"tokens_out":1063,"duration_ms":10453,"temperature":1.0,"reasoning_tokens":988,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:51:52.415853+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human experts blind-review the same 28 papers, along with human-written accepted papers, and compare their scores and defect categories with DeepReviewer-14B's; if humans rate the AI papers near parity with accepted human work, or if DeepReviewer-14B gives equally low scores to strong human papers, the central claim is unsupported.","supporting_citations":[],"review_version":1}