REVIEW 3 major objections 5 minor 19 references
An agentic pipeline grounds every claim in real literature, runs actual bioinformatics experiments, and lifts manuscript quality by nearly 18 points through deep research cycles.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 18:14 UTC pith:PAT2R43F
load-bearing objection Useful integrated research-assistant pipeline with real experiments and clean citation bounds; the +17.96 “quality” gain is partly self-scored and should be read as an internal trajectory, not publishability proof. the 3 major comments →
Prompt-to-Paper: Agentic AI System for Bioinformatics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A multi-agent pipeline that binds every claim to a retrieved literature corpus, injects only numbers from executed bioinformatics experiments, and drives revision with an eight-dimensional scorer plus periodic deep research cycles can produce submission-formatted manuscripts that improve by an average of +17.96 points (maximum +26.04) on a 0–100 scale across five case studies, at roughly $0.31 per paper and with zero out-of-range citations.
What carries the argument
The quality-driven improvement loop with deep research cycles: each iteration routes the weakest of eight scored dimensions to one of three actions (add statistical analysis from executed results, gather new literature evidence, or rewrite for clarity), and every ten iterations the system identifies a scientific gap, extends and re-runs the experiment, then re-manuscripts the full paper from the stronger outputs under a never-regress acceptance rule.
Load-bearing premise
The reported quality gains assume that climbing an internal hybrid scorer—heavily weighted toward the same model family used to write and judge the papers—tracks genuine scientific quality rather than optimization to that scorer.
What would settle it
Have independent domain experts score the before- and after-improvement manuscripts on the same eight dimensions while blinded to condition; if the after drafts show no reliable gain, or if re-running the released experiment scripts outside the pipeline fails to recover the reported numbers, the central quality claim does not hold.
If this is right
- Bioinformatics first drafts can be generated with zero fabricated citations when every claim is bound to a retrieved corpus of real papers.
- Injecting only executed experiment numbers via a canonical results file removes a major source of numeric hallucination.
- Deep research cycles that re-run experiments produce larger score jumps than prose polishing alone.
- At roughly thirty-one cents per paper, large-scale generation of grounded first drafts becomes practical.
- Presentation and structural completeness remain the primary dimensions still requiring human revision before submission.
Where Pith is reading between the lines
- The same loop could transfer to other computational sciences if the coding agent is given domain-appropriate tools and reference data.
- A controlled human-expert calibration study on matched topics would be the direct next step to test whether the automated score tracks publishability.
- Relaxing the strict never-regress rule to allow small overall dips when the targeted dimension jumps substantially might extract more gain per iteration budget.
- Self-referential corpus metrics and a generation-family judge may under-reward genuine novelty that lies outside the retrieved seed set.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Prompt-to-Paper (RLEv4), a multi-agent pipeline that turns a bioinformatics topic into a submission-formatted manuscript. It combines (i) deterministic RAG over a 60–100 paper corpus with section-aware scoring and snowball expansion, (ii) an autonomous coding agent that runs real computational experiments and injects verified numbers via a CanonicalResults object, and (iii) an eight-dimensional hybrid quality scorer (G-Eval + entity-network + heuristics, with hallucination penalties) driving a context-rich improvement loop with deep research cycles every ten iterations. On five bioinformatics case studies the system reports compiled PDFs, zero out-of-range citations, mean quality gain +17.96/100 over 60 iterations (max +26.04), independent LLM reviewer average ~7.39/10, a single human average 7.0/10, and cost ≈ $0.31 per paper.
Significance. If the operational claims hold, the work is a useful systems contribution to automated scientific writing: it replaces fabricated experimental numbers with executed outputs, enforces corpus-bound citations, and couples revision to an explicit multi-dimensional score with deep re-experimentation cycles. Open code, detailed Algorithm 1, token/cost accounting, and honest limitation statements are strengths. The advance is primarily engineering and evaluation methodology for agentic paper generation in computational biology, not a new scientific result in bioinformatics itself. Significance for the field depends on whether reported score gains track independent quality rather than scorer-local optimization.
major comments (3)
- [§III-E, §III-F, Table I] §III-E/F, Algorithm 1, Table I: The central +17.96 claim is measured by a hybrid scorer whose Tier-1 component (55% weight) is deepseek-v4-pro G-Eval—the same leader model used for gap-finding, re-manuscripting, and deep-cycle rewriting (§III-C). The never-regress rule and stabilized re-scoring therefore optimize against a judge that shares the generator’s model family and role stack. Approximate ICLR z-normalization and self-referential Tier-2 metrics do not break this coupling. External LLM/human scores assess only final drafts, not the before→after delta. Without an ablation that freezes or swaps the judge (e.g., held-out model family, or human before/after on the same five drafts), the trajectory cannot be taken as scorer-independent scientific improvement.
- [§IV-D, §IV-F, Table IV, Table VII] §IV-D Table IV vs §IV-F Table VII / Appendix A: The paper asserts zero bad citations (every [N] within corpus bounds) while the human reviewer repeatedly flags “References fake,” empty brackets, and inconsistent reference formatting on the reproduced manuscripts. These are not the same failure modes. The manuscript should define citation fidelity more carefully (range check vs. support for the claim vs. bibliographic correctness) and report how many human-flagged citation defects remain after the hallucination audit. As written, the strong “zero out-of-range” claim and the human annotations sit in unresolved tension.
- [§IV, Limitations, Abstract] §IV-A–G, Limitations: Evaluation rests on n=5 bioinformatics problems, self-referential Recall@20, non-controlled literature comparisons (Table VIII), and a single human rater. The paper acknowledges several of these points, but the abstract and conclusion still present +17.96, 7.0/10, and “measurably good” as general evidence of publishability-oriented quality. Either expand the evaluation (more problems, blinded multi-expert before/after, or a fixed external judge) or substantially narrow the claim language so that results are framed as internal pipeline metrics plus partial external checks on five drafts.
minor comments (5)
- [Table II, §IV-B] Table II: structural completeness averages 46.84 and is highly variable (32.50–67.65); presentation averages 57.58. These weak dimensions should be foregrounded earlier when interpreting the B–B+ band, not only in the summary.
- [§III-E] Figure 4 and Eq. for Ph: the hallucination penalty coefficients (0.5, 0.3, 0.2) and the global 0.3 factor on Q are free parameters; a short sensitivity note would help readers judge robustness.
- [§III-A] Eq. (1): section weights ws are “empirically set”; state the pilot procedure or held-out topic set used to choose them, even briefly.
- [Appendix A, §IV-F] Appendix A manuscripts are valuable as unedited artefacts, but several human notes (missing figures, unstructured abstracts, grammar) should be cross-referenced in the main text when claiming “submission-formatted” PDFs so readers do not over-read formatting success as content readiness.
- [§IV-G, Table VIII] Related work comparisons (CycleResearcher, AI Scientist, EpidemIQs) are useful context; keep the explicit non-comparability caveat of Table VIII adjacent to any cost or score juxtaposition in the prose, not only in the table note.
Circularity Check
The +17.96 quality gain is produced by an improvement loop that accepts revisions only when a hybrid scorer rises, and that scorer’s dominant (55%) G-Eval tier is the identical deepseek-v4-pro leader used for generation, gap-finding and re-manuscripting.
specific steps
-
fitted input called prediction
[§III-C, §III-E (Tier 1), Algorithm 1 lines 15–27, §IV-A Table I]
"The leader role (quality judging, gap-finding, synthesis, and deep-cycle re-manuscripting) uses deepseek-v4-pro … Tier 1 (55%): G-Eval LLM-as-judge. The leader model (deepseek-v4-pro) scores each dimension … accepts a candidate revision only when it improves the overall score … The improvement loop raises manuscript quality by an average of +17.96 points"
The quantity reported as the principal result (+17.96) is exactly the objective that the loop is constructed to maximize; the dominant judge is the identical model family and role that generates and rewrites the manuscript. Score gains are therefore the outcome of optimizing against that judge by construction, not an independent external measurement of scientific quality.
-
self definitional
[§III-E (hallucination audit), Table IV, §IV-D]
"The citation-range check uses the strict bound max_valid_ref = len(corpus), ensuring any [N] marker outside the actual retrieved corpus is flagged. … The system achieves zero bad citations across all five manuscripts: every [N] marker … corresponds to a paper that genuinely exists in the retrieved corpus."
Zero out-of-range citations is enforced by the deterministic bound that defines the valid reference set as precisely the papers the pipeline itself retrieved; the reported fidelity metric is therefore true by construction of the citation mechanism rather than an independent empirical discovery.
-
self definitional
[§IV-D Table IV note, §IV-I Limitations]
"Recall@20 = fraction of the top-20 knowledge-graph papers cited in the manuscript; self-referential, measured against the same corpus used for retrieval. … Recall@20 is self-referential. It is measured against the same knowledge graph used for retrieval, so it reflects internal citation consistency rather than external completeness."
The metric is defined against the identical graph that the retrieval stage produced; high Recall@20 therefore restates internal consistency of the system’s own corpus rather than measuring external grounding.
full rationale
The paper’s central empirical claim is that a 60-iteration quality-driven loop raises an eight-dimensional score by +17.96 (max +26.04) on five bioinformatics manuscripts. That score is defined by a three-tier hybrid whose Tier-1 component (weight 0.55) is G-Eval performed by deepseek-v4-pro—the same model that serves as leader for planning, gap-finding, deep-cycle re-manuscripting and section synthesis. Acceptance is strictly never-regress on this score (or a targeted dimension under a tie). Consequently the reported trajectory is the result of optimizing the very objective that later measures success; stabilized re-judging of only changed dimensions and deterministic hallucination penalties reduce noise but do not open the loop. Zero out-of-range citations is likewise guaranteed by the hard bound max_valid_ref = len(corpus) together with the pipeline’s exclusive use of that corpus. External LLM and single-human averages (~7/10) supply a partial consistency check on absolute level but do not validate that the internal delta itself is scorer-independent. The engineering artefacts (real code execution, corpus-bound citations, compiled PDFs) remain non-circular; the quality-gain claim, however, partially reduces by construction to the closed generator–judge pair. This is partial rather than total circularity, hence score 6.
Axiom & Free-Parameter Ledger
free parameters (6)
- Section relevance weights w_s =
0.30/0.25/0.20/0.12/0.08/0.05
- Snowball relevance threshold τ =
τ ∈ [0.05, 0.08]
- Hybrid scorer tier weights =
0.55 / 0.25 / 0.20
- ICLR-approximate dimension calibrations (μ_d, σ_d) =
approximate ICLR 2023–2025 per-dimension μ,σ
- IMPROVE_ITERATIONS and deep-cycle period K =
60 iterations; K=10
- Hallucination penalty coefficients =
0.5 / 0.3 / 0.2; global ×0.3
axioms (5)
- domain assumption A retrieved Semantic Scholar/Tavily corpus of 60–100 papers plus strict [N] range checks is sufficient to deterministically ground manuscript claims.
- domain assumption Hardcoded sequences/matrices with local Python execution constitute 'real computational biology experiments' whose numbers validate scientific claims in the draft.
- ad hoc to paper Approximate ICLR score calibrations and deepseek-v4-pro G-Eval judgments are valid enough proxies for multi-dimensional manuscript quality.
- ad hoc to paper Never-regress acceptance on overall Q (or targeted dimension on ties) yields genuine quality improvement rather than only scorer-local optima.
- standard math Standard embedding similarity (SPECTER2/BM25) and PageRank/claim-relation LLM labels adequately represent literature relevance and contradiction structure.
invented entities (4)
-
RLEv4 / Prompt-to-Paper multi-agent pipeline
independent evidence
-
Eight-dimensional hybrid quality scorer with hallucination audit
no independent evidence
-
ContextRichImprover with deep research cycles
independent evidence
-
CanonicalResults object injected into every section
independent evidence
read the original abstract
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication. We present Prompt-to-Paper, a multi-agent framework that directly addresses this evaluation gap through three integrated innovations. First, a deterministic retrieval-augmented generation pipeline with section-aware relevance scoring and snowball citation expansion grounds every claim in a verifiable corpus of 60--100 papers. Second, an autonomous coding agent executes real computational biology experiments replacing synthetic outputs with genuine numerical results. Third, an eight-dimensional automated quality scorer, benchmarked with approximate reference statistics from published papers and augmented with explicit hallucination penalties, provides standardized, reproducible quality assessments. The quality-driven improvement loop uses a context-rich reviser that routes each iteration to one of three researcher actions and fires a deep research cycle every ten iterations to re-run experiments and re-manuscript from stronger outputs. We validate the system on five bioinformatics case studies; all five cases compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raises manuscript quality by an average of +17.96 points on a 0--100 scale (maximum +26.04. As partial external checks, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Complete manuscripts are produced at approximately 0.31 USD per paper.
Figures
Reference graph
Works this paper leans on
-
[1]
Scientific literature: Information overload,
E. Landhuis, “Scientific literature: Information overload,”Nature, vol. 535, no. 7612, pp. 457–458, 2016. [Online]. Available: https://doi.org/10.1038/nj7612-457a
-
[2]
Why most published research findings are false,
J. P. A. Ioannidis, “Why most published research findings are false,” PLoS Medicine, vol. 2, no. 8, p. e124, 2005. [Online]. Available: https://doi.org/10.1371/journal.pmed.0020124
-
[3]
1,500 scientists lift the lid on reproducibility,
M. Baker, “1,500 scientists lift the lid on reproducibility,”Nature, vol. 533, no. 7604, pp. 452–454, 2016. [Online]. Available: https://doi.org/10.1038/533452a
doi:10.1038/533452a 2016
-
[4]
J. Brainard, “Rethinking retractions,”Science, vol. 362, no. 6413, pp. 390–393, 2018. [Online]. Available: https://doi.org/10.1126/science. 362.6413.390
doi:10.1126/science 2018
-
[5]
R. M. Kwee and T. C. Kwee, “Retracted publications in medical imaging literature: an analysis using the retraction watch database,” Academic Radiology, vol. 30, no. 6, pp. 1148–1152, 2023. [Online]. Available: https://doi.org/10.1016/j.acra.2022.06.025
-
[6]
RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway,
X. Guo, L. Dong, and D. Hao, “RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway,” Frontiers in Cell and Developmental Biology, vol. 11, p. 1339390,
-
[7]
Available: https://doi.org/10.3389/fcell.2023.1339390
[Online]. Available: https://doi.org/10.3389/fcell.2023.1339390
-
[9]
Available: https://arxiv.org/abs/2504.08066
[Online]. Available: https://arxiv.org/abs/2504.08066
-
[10]
Jr. AI scientist and its risk report: Autonomous scientific exploration from a baseline paper,
A. Miyai, M. Toyooka, T. Otonari, Z. Zhao, and K. Aizawa, “Jr. AI scientist and its risk report: Autonomous scientific exploration from a baseline paper,”arXiv preprint arXiv:2511.04583, 2025. [Online]. Available: https://arxiv.org/abs/2511.04583
arXiv 2025
-
[11]
Towards a transparent and reproducible AI-assisted research paper writing,
J. Park, “Towards a transparent and reproducible AI-assisted research paper writing,”Genomics & Informatics, vol. 23, no. 1, p. 26, 2025. [Online]. Available: https://doi.org/10.1186/s44342-025-00057-0
-
[12]
PaperRobot: Incremental draft generation of scientific ideas,
Q. Wang, L. Huang, Z. Jiang, K. Knight, H. Ji, M. Bansal, and Y . Luan, “PaperRobot: Incremental draft generation of scientific ideas,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 1980–1991. [Online]. Available: https://doi.org/10.18653/v1/P19-1191
-
[13]
PaSa: An LLM agent for comprehensive academic paper search,
Y . He, G. Huang, P. Feng, Y . Lin, Y . b. Zhang, and H. Li, “PaSa: An LLM agent for comprehensive academic paper search,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 11 663–11 679. [Online]. Available: https://arxiv.org/abs/2501.10120
Pith/arXiv arXiv 2025
-
[14]
ScholarGym: Benchmarking deep research workflows on academic literature retrieval,
H. Shen, H. Yang, and Z. Gu, “ScholarGym: Benchmarking deep research workflows on academic literature retrieval,”arXiv preprint arXiv:2601.21654, 2026. [Online]. Available: https://arxiv.org/abs/2601. 21654
arXiv 2026
-
[15]
CycleResearcher: Improving automated research via automated review,
Y . Weng, M. Zhu, G. Bao, H. Zhang, J. Wang, Y . Zhang, and L. Yang, “CycleResearcher: Improving automated research via automated review,” inInternational Conference on Learning Representations,
-
[16]
Available: https://arxiv.org/abs/2411.00816
[Online]. Available: https://arxiv.org/abs/2411.00816
-
[17]
OpenLens AI: Fully autonomous research agent for health informatics,
Y . Cheng and J. Suo, “OpenLens AI: Fully autonomous research agent for health informatics,”arXiv preprint arXiv:2509.14778, 2025. [Online]. Available: https://arxiv.org/abs/2509.14778
arXiv 2025
-
[18]
Accelerating scientific research with Gemini: Case studies and common techniques,
D. P. Woodruff, V . Cohen-Addad, L. Jain, J. Mao, S. Zuo, M. Bateni, S. b. Branzei, M. P. Brenner, L. Chen, and Y . Feng, “Accelerating scientific research with Gemini: Case studies and common techniques,”arXiv preprint arXiv:2602.03837, 2026. [Online]. Available: https://arxiv.org/abs/2602.03837
arXiv 2026
-
[19]
EpidemIQs: Prompt-to-paper LLM agents for epidemic modeling and analysis,
M. H. Samaei, F. D. Sahneh, L. W. Cohnstaedt, and C. M. Scoglio, “EpidemIQs: Prompt-to-paper LLM agents for epidemic modeling and analysis,”arXiv preprint arXiv:2510.00024, 2025. [Online]. Available: https://arxiv.org/abs/2510.00024
arXiv 2025
-
[20]
NORA: A harness-engineered autonomous research agent for end-to-end spatial data science,
B. Zhou, X. Huang, H. Ning, Q. Wu, D. Li, and Z. Zhang, “NORA: A harness-engineered autonomous research agent for end-to-end spatial data science,”arXiv preprint arXiv:2605.02092, 2026. [Online]. Available: https://arxiv.org/abs/2605.02092 APPENDIXA GENERATEDMANUSCRIPTS: UPDATEDPIPELINE The following pages reproduce the complete RLEv4- generated manuscrip...
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.