Pith. sign in

REVIEW 3 major objections 5 minor 19 references

An agentic pipeline grounds every claim in real literature, runs actual bioinformatics experiments, and lifts manuscript quality by nearly 18 points through deep research cycles.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 18:14 UTC pith:PAT2R43F

load-bearing objection Useful integrated research-assistant pipeline with real experiments and clean citation bounds; the +17.96 “quality” gain is partly self-scored and should be read as an internal trajectory, not publishability proof. the 3 major comments →

arxiv 2607.05456 v1 pith:PAT2R43F submitted 2026-07-05 cs.AI cs.CLq-bio.QM

Prompt-to-Paper: Agentic AI System for Bioinformatics

classification cs.AI cs.CLq-bio.QM
keywords agentic AIbioinformaticsautomated manuscript generationretrieval-augmented generationscientific quality assessmentdeep research cyclesautonomous coding agenthallucination penalties
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Automated paper writers often invent citations and fake experimental numbers, and they have no standard way to check whether the result is fit for real publication. This paper offers a multi-agent system that first retrieves a verifiable corpus of 60–100 real papers, then runs genuine computational biology experiments whose numbers alone appear in the text, and finally scores each draft on eight dimensions with explicit penalties for hallucinations. A context-rich loop repeatedly targets the weakest dimension—adding analysis, gathering evidence, or rewriting—and every ten iterations launches a deep research cycle that re-runs experiments and rewrites the whole manuscript from stronger results. On five bioinformatics case studies the system produced compiled submission-formatted PDFs with zero out-of-range citations, raised the automated quality score by an average of 17.96 points, and cost about thirty-one cents per paper. External checks by other language models and a human reviewer placed the drafts near seven out of ten, framing the system as a source of competent first drafts rather than finished publications.

Core claim

A multi-agent pipeline that binds every claim to a retrieved literature corpus, injects only numbers from executed bioinformatics experiments, and drives revision with an eight-dimensional scorer plus periodic deep research cycles can produce submission-formatted manuscripts that improve by an average of +17.96 points (maximum +26.04) on a 0–100 scale across five case studies, at roughly $0.31 per paper and with zero out-of-range citations.

What carries the argument

The quality-driven improvement loop with deep research cycles: each iteration routes the weakest of eight scored dimensions to one of three actions (add statistical analysis from executed results, gather new literature evidence, or rewrite for clarity), and every ten iterations the system identifies a scientific gap, extends and re-runs the experiment, then re-manuscripts the full paper from the stronger outputs under a never-regress acceptance rule.

Load-bearing premise

The reported quality gains assume that climbing an internal hybrid scorer—heavily weighted toward the same model family used to write and judge the papers—tracks genuine scientific quality rather than optimization to that scorer.

What would settle it

Have independent domain experts score the before- and after-improvement manuscripts on the same eight dimensions while blinded to condition; if the after drafts show no reliable gain, or if re-running the released experiment scripts outside the pipeline fails to recover the reported numbers, the central quality claim does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Bioinformatics first drafts can be generated with zero fabricated citations when every claim is bound to a retrieved corpus of real papers.
  • Injecting only executed experiment numbers via a canonical results file removes a major source of numeric hallucination.
  • Deep research cycles that re-run experiments produce larger score jumps than prose polishing alone.
  • At roughly thirty-one cents per paper, large-scale generation of grounded first drafts becomes practical.
  • Presentation and structural completeness remain the primary dimensions still requiring human revision before submission.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same loop could transfer to other computational sciences if the coding agent is given domain-appropriate tools and reference data.
  • A controlled human-expert calibration study on matched topics would be the direct next step to test whether the automated score tracks publishability.
  • Relaxing the strict never-regress rule to allow small overall dips when the targeted dimension jumps substantially might extract more gain per iteration budget.
  • Self-referential corpus metrics and a generation-family judge may under-reward genuine novelty that lies outside the retrieved seed set.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Prompt-to-Paper (RLEv4), a multi-agent pipeline that turns a bioinformatics topic into a submission-formatted manuscript. It combines (i) deterministic RAG over a 60–100 paper corpus with section-aware scoring and snowball expansion, (ii) an autonomous coding agent that runs real computational experiments and injects verified numbers via a CanonicalResults object, and (iii) an eight-dimensional hybrid quality scorer (G-Eval + entity-network + heuristics, with hallucination penalties) driving a context-rich improvement loop with deep research cycles every ten iterations. On five bioinformatics case studies the system reports compiled PDFs, zero out-of-range citations, mean quality gain +17.96/100 over 60 iterations (max +26.04), independent LLM reviewer average ~7.39/10, a single human average 7.0/10, and cost ≈ $0.31 per paper.

Significance. If the operational claims hold, the work is a useful systems contribution to automated scientific writing: it replaces fabricated experimental numbers with executed outputs, enforces corpus-bound citations, and couples revision to an explicit multi-dimensional score with deep re-experimentation cycles. Open code, detailed Algorithm 1, token/cost accounting, and honest limitation statements are strengths. The advance is primarily engineering and evaluation methodology for agentic paper generation in computational biology, not a new scientific result in bioinformatics itself. Significance for the field depends on whether reported score gains track independent quality rather than scorer-local optimization.

major comments (3)
  1. [§III-E, §III-F, Table I] §III-E/F, Algorithm 1, Table I: The central +17.96 claim is measured by a hybrid scorer whose Tier-1 component (55% weight) is deepseek-v4-pro G-Eval—the same leader model used for gap-finding, re-manuscripting, and deep-cycle rewriting (§III-C). The never-regress rule and stabilized re-scoring therefore optimize against a judge that shares the generator’s model family and role stack. Approximate ICLR z-normalization and self-referential Tier-2 metrics do not break this coupling. External LLM/human scores assess only final drafts, not the before→after delta. Without an ablation that freezes or swaps the judge (e.g., held-out model family, or human before/after on the same five drafts), the trajectory cannot be taken as scorer-independent scientific improvement.
  2. [§IV-D, §IV-F, Table IV, Table VII] §IV-D Table IV vs §IV-F Table VII / Appendix A: The paper asserts zero bad citations (every [N] within corpus bounds) while the human reviewer repeatedly flags “References fake,” empty brackets, and inconsistent reference formatting on the reproduced manuscripts. These are not the same failure modes. The manuscript should define citation fidelity more carefully (range check vs. support for the claim vs. bibliographic correctness) and report how many human-flagged citation defects remain after the hallucination audit. As written, the strong “zero out-of-range” claim and the human annotations sit in unresolved tension.
  3. [§IV, Limitations, Abstract] §IV-A–G, Limitations: Evaluation rests on n=5 bioinformatics problems, self-referential Recall@20, non-controlled literature comparisons (Table VIII), and a single human rater. The paper acknowledges several of these points, but the abstract and conclusion still present +17.96, 7.0/10, and “measurably good” as general evidence of publishability-oriented quality. Either expand the evaluation (more problems, blinded multi-expert before/after, or a fixed external judge) or substantially narrow the claim language so that results are framed as internal pipeline metrics plus partial external checks on five drafts.
minor comments (5)
  1. [Table II, §IV-B] Table II: structural completeness averages 46.84 and is highly variable (32.50–67.65); presentation averages 57.58. These weak dimensions should be foregrounded earlier when interpreting the B–B+ band, not only in the summary.
  2. [§III-E] Figure 4 and Eq. for Ph: the hallucination penalty coefficients (0.5, 0.3, 0.2) and the global 0.3 factor on Q are free parameters; a short sensitivity note would help readers judge robustness.
  3. [§III-A] Eq. (1): section weights ws are “empirically set”; state the pilot procedure or held-out topic set used to choose them, even briefly.
  4. [Appendix A, §IV-F] Appendix A manuscripts are valuable as unedited artefacts, but several human notes (missing figures, unstructured abstracts, grammar) should be cross-referenced in the main text when claiming “submission-formatted” PDFs so readers do not over-read formatting success as content readiness.
  5. [§IV-G, Table VIII] Related work comparisons (CycleResearcher, AI Scientist, EpidemIQs) are useful context; keep the explicit non-comparability caveat of Table VIII adjacent to any cost or score juxtaposition in the prose, not only in the table note.

Circularity Check

3 steps flagged

The +17.96 quality gain is produced by an improvement loop that accepts revisions only when a hybrid scorer rises, and that scorer’s dominant (55%) G-Eval tier is the identical deepseek-v4-pro leader used for generation, gap-finding and re-manuscripting.

specific steps
  1. fitted input called prediction [§III-C, §III-E (Tier 1), Algorithm 1 lines 15–27, §IV-A Table I]
    "The leader role (quality judging, gap-finding, synthesis, and deep-cycle re-manuscripting) uses deepseek-v4-pro … Tier 1 (55%): G-Eval LLM-as-judge. The leader model (deepseek-v4-pro) scores each dimension … accepts a candidate revision only when it improves the overall score … The improvement loop raises manuscript quality by an average of +17.96 points"

    The quantity reported as the principal result (+17.96) is exactly the objective that the loop is constructed to maximize; the dominant judge is the identical model family and role that generates and rewrites the manuscript. Score gains are therefore the outcome of optimizing against that judge by construction, not an independent external measurement of scientific quality.

  2. self definitional [§III-E (hallucination audit), Table IV, §IV-D]
    "The citation-range check uses the strict bound max_valid_ref = len(corpus), ensuring any [N] marker outside the actual retrieved corpus is flagged. … The system achieves zero bad citations across all five manuscripts: every [N] marker … corresponds to a paper that genuinely exists in the retrieved corpus."

    Zero out-of-range citations is enforced by the deterministic bound that defines the valid reference set as precisely the papers the pipeline itself retrieved; the reported fidelity metric is therefore true by construction of the citation mechanism rather than an independent empirical discovery.

  3. self definitional [§IV-D Table IV note, §IV-I Limitations]
    "Recall@20 = fraction of the top-20 knowledge-graph papers cited in the manuscript; self-referential, measured against the same corpus used for retrieval. … Recall@20 is self-referential. It is measured against the same knowledge graph used for retrieval, so it reflects internal citation consistency rather than external completeness."

    The metric is defined against the identical graph that the retrieval stage produced; high Recall@20 therefore restates internal consistency of the system’s own corpus rather than measuring external grounding.

full rationale

The paper’s central empirical claim is that a 60-iteration quality-driven loop raises an eight-dimensional score by +17.96 (max +26.04) on five bioinformatics manuscripts. That score is defined by a three-tier hybrid whose Tier-1 component (weight 0.55) is G-Eval performed by deepseek-v4-pro—the same model that serves as leader for planning, gap-finding, deep-cycle re-manuscripting and section synthesis. Acceptance is strictly never-regress on this score (or a targeted dimension under a tie). Consequently the reported trajectory is the result of optimizing the very objective that later measures success; stabilized re-judging of only changed dimensions and deterministic hallucination penalties reduce noise but do not open the loop. Zero out-of-range citations is likewise guaranteed by the hard bound max_valid_ref = len(corpus) together with the pipeline’s exclusive use of that corpus. External LLM and single-human averages (~7/10) supply a partial consistency check on absolute level but do not validate that the internal delta itself is scorer-independent. The engineering artefacts (real code execution, corpus-bound citations, compiled PDFs) remain non-circular; the quality-gain claim, however, partially reduces by construction to the closed generator–judge pair. This is partial rather than total circularity, hence score 6.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 4 invented entities

The central performance claims rest on engineering choices and evaluation axioms rather than free physical constants. Load-bearing free parameters include scorer tier weights, section relevance weights, relevance threshold τ, ICLR-style calibration means/stds, iteration budget, and deep-cycle cadence. Domain assumptions include that Semantic Scholar retrieval plus SPECTER2/BM25 scoring yields a sufficient verifiable corpus, that hardcoded computational-biology tasks stand in for real experimental science, and that the hybrid scorer approximates publication quality. The main invented entities are the RLEv4 pipeline components themselves (ContextRichImprover, CanonicalResults injection, eight-dimension hybrid scorer), which are software constructs with independent evidence only insofar as the released code and reported runs can be re-executed.

free parameters (6)
  • Section relevance weights w_s = 0.30/0.25/0.20/0.12/0.08/0.05
    Empirically set weights (abstract 0.30, methods 0.25, results 0.20, introduction 0.12, discussion 0.08, conclusion 0.05) control which papers enter the grounded corpus and thus every later claim.
  • Snowball relevance threshold τ = τ ∈ [0.05, 0.08]
    Threshold band [0.05, 0.08] decides which cited/citing papers expand the 60–100 paper corpus.
  • Hybrid scorer tier weights = 0.55 / 0.25 / 0.20
    0.55 G-Eval / 0.25 entity-network / 0.20 heuristics dominate the overall Q that the improvement loop optimizes.
  • ICLR-approximate dimension calibrations (μ_d, σ_d) = approximate ICLR 2023–2025 per-dimension μ,σ
    Used to z-normalize 0–10 G-Eval scores into 0–100; approximate reference statistics, not a controlled human study.
  • IMPROVE_ITERATIONS and deep-cycle period K = 60 iterations; K=10
    Default 60 iterations with deep cycles every 10 iterations define the measured +17.96 gain regime.
  • Hallucination penalty coefficients = 0.5 / 0.3 / 0.2; global ×0.3
    Ph formula coefficients (0.5 cite, 0.3 numeric mismatch, 0.2 coverage; global 0.3) directly alter grounding/soundness and overall Q.
axioms (5)
  • domain assumption A retrieved Semantic Scholar/Tavily corpus of 60–100 papers plus strict [N] range checks is sufficient to deterministically ground manuscript claims.
    Stated as innovation (i) and enforced in the citation-range audit; does not guarantee semantic correctness of citations.
  • domain assumption Hardcoded sequences/matrices with local Python execution constitute 'real computational biology experiments' whose numbers validate scientific claims in the draft.
    §III-D autonomous coding agent; no network calls during experiments except optional NCBI fallbacks.
  • ad hoc to paper Approximate ICLR score calibrations and deepseek-v4-pro G-Eval judgments are valid enough proxies for multi-dimensional manuscript quality.
    §III-E three-tier scorer; paper later admits lack of controlled human calibration.
  • ad hoc to paper Never-regress acceptance on overall Q (or targeted dimension on ties) yields genuine quality improvement rather than only scorer-local optima.
    §III-F improvement loop acceptance rule and deep-cycle adoption guard.
  • standard math Standard embedding similarity (SPECTER2/BM25) and PageRank/claim-relation LLM labels adequately represent literature relevance and contradiction structure.
    Eq. (1) and knowledge-graph construction in §III-A/B.
invented entities (4)
  • RLEv4 / Prompt-to-Paper multi-agent pipeline independent evidence
    purpose: End-to-end conversion of a topic into a scored, PDF-compiled bioinformatics manuscript with real experiment injection.
    Primary system contribution; evidence is the reported runs and claimed open-source release, not an external physical entity.
  • Eight-dimensional hybrid quality scorer with hallucination audit no independent evidence
    purpose: Provide standardized 0–100 quality assessments that drive the revision loop.
    New composite instrument for this paper; independent evidence limited to internal consistency and partial external LLM/human checks.
  • ContextRichImprover with deep research cycles independent evidence
    purpose: Route revisions to ADD_ANALYSIS / GATHER_EVIDENCE / REWRITE and periodically re-run experiments then re-manuscript.
    Software control loop presented as the mechanism behind sustained score gains past prose polishing.
  • CanonicalResults object injected into every section independent evidence
    purpose: Guarantee numerical consistency between executed experiments and manuscript text.
    Internal data structure; falsifiable only by re-running the agent and checking results.json vs text.

pith-pipeline@v1.1.0-grok45 · 20420 in / 4447 out tokens · 46287 ms · 2026-07-11T18:14:06.197291+00:00 · methodology

0 comments
read the original abstract

While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication. We present Prompt-to-Paper, a multi-agent framework that directly addresses this evaluation gap through three integrated innovations. First, a deterministic retrieval-augmented generation pipeline with section-aware relevance scoring and snowball citation expansion grounds every claim in a verifiable corpus of 60--100 papers. Second, an autonomous coding agent executes real computational biology experiments replacing synthetic outputs with genuine numerical results. Third, an eight-dimensional automated quality scorer, benchmarked with approximate reference statistics from published papers and augmented with explicit hallucination penalties, provides standardized, reproducible quality assessments. The quality-driven improvement loop uses a context-rich reviser that routes each iteration to one of three researcher actions and fires a deep research cycle every ten iterations to re-run experiments and re-manuscript from stronger outputs. We validate the system on five bioinformatics case studies; all five cases compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raises manuscript quality by an average of +17.96 points on a 0--100 scale (maximum +26.04. As partial external checks, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Complete manuscripts are produced at approximately 0.31 USD per paper.

Figures

Figures reproduced from arXiv: 2607.05456 by Arsalan Shaukat, Maheera Amjad, Muhammad U.S. Khan, Ramsha Kamran, Salma Sherbaz, Zartasha Mustansar.

Figure 1
Figure 1. Figure 1: Planning-agent output for the TP53 hotspot mutation [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Knowledge graph of claim alignments for the TP53 query. Nodes are papers, sized by PageRank; edges are coloured [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Papers view of the RLEv4 interactive dashboard for the TP53 hotspot-mutation query, showing corpus statistics, [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: RLEv4 evaluation architecture. The three-tier hybrid [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: RLEv4 improvement loop architecture. Every [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Improvement loop convergence for the Substitution Matrix Eigenspectrum case (representative of all five runs). [PITH_FULL_IMAGE:figures/full_fig_p061_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

19 extracted references · 4 canonical work pages

  1. [1]

    Scientific literature: Information overload,

    E. Landhuis, “Scientific literature: Information overload,”Nature, vol. 535, no. 7612, pp. 457–458, 2016. [Online]. Available: https://doi.org/10.1038/nj7612-457a

  2. [2]

    Why most published research findings are false,

    J. P. A. Ioannidis, “Why most published research findings are false,” PLoS Medicine, vol. 2, no. 8, p. e124, 2005. [Online]. Available: https://doi.org/10.1371/journal.pmed.0020124

  3. [3]

    1,500 scientists lift the lid on reproducibility,

    M. Baker, “1,500 scientists lift the lid on reproducibility,”Nature, vol. 533, no. 7604, pp. 452–454, 2016. [Online]. Available: https://doi.org/10.1038/533452a

  4. [4]

    Rethinking retractions,

    J. Brainard, “Rethinking retractions,”Science, vol. 362, no. 6413, pp. 390–393, 2018. [Online]. Available: https://doi.org/10.1126/science. 362.6413.390

  5. [5]

    Retracted publications in medical imaging literature: an analysis using the retraction watch database,

    R. M. Kwee and T. C. Kwee, “Retracted publications in medical imaging literature: an analysis using the retraction watch database,” Academic Radiology, vol. 30, no. 6, pp. 1148–1152, 2023. [Online]. Available: https://doi.org/10.1016/j.acra.2022.06.025

  6. [6]

    RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway,

    X. Guo, L. Dong, and D. Hao, “RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway,” Frontiers in Cell and Developmental Biology, vol. 11, p. 1339390,

  7. [7]

    Available: https://doi.org/10.3389/fcell.2023.1339390

    [Online]. Available: https://doi.org/10.3389/fcell.2023.1339390

  8. [9]

    Available: https://arxiv.org/abs/2504.08066

    [Online]. Available: https://arxiv.org/abs/2504.08066

  9. [10]

    Jr. AI scientist and its risk report: Autonomous scientific exploration from a baseline paper,

    A. Miyai, M. Toyooka, T. Otonari, Z. Zhao, and K. Aizawa, “Jr. AI scientist and its risk report: Autonomous scientific exploration from a baseline paper,”arXiv preprint arXiv:2511.04583, 2025. [Online]. Available: https://arxiv.org/abs/2511.04583

  10. [11]

    Towards a transparent and reproducible AI-assisted research paper writing,

    J. Park, “Towards a transparent and reproducible AI-assisted research paper writing,”Genomics & Informatics, vol. 23, no. 1, p. 26, 2025. [Online]. Available: https://doi.org/10.1186/s44342-025-00057-0

  11. [12]

    PaperRobot: Incremental draft generation of scientific ideas,

    Q. Wang, L. Huang, Z. Jiang, K. Knight, H. Ji, M. Bansal, and Y . Luan, “PaperRobot: Incremental draft generation of scientific ideas,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 1980–1991. [Online]. Available: https://doi.org/10.18653/v1/P19-1191

  12. [13]

    PaSa: An LLM agent for comprehensive academic paper search,

    Y . He, G. Huang, P. Feng, Y . Lin, Y . b. Zhang, and H. Li, “PaSa: An LLM agent for comprehensive academic paper search,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 11 663–11 679. [Online]. Available: https://arxiv.org/abs/2501.10120

  13. [14]

    ScholarGym: Benchmarking deep research workflows on academic literature retrieval,

    H. Shen, H. Yang, and Z. Gu, “ScholarGym: Benchmarking deep research workflows on academic literature retrieval,”arXiv preprint arXiv:2601.21654, 2026. [Online]. Available: https://arxiv.org/abs/2601. 21654

  14. [15]

    CycleResearcher: Improving automated research via automated review,

    Y . Weng, M. Zhu, G. Bao, H. Zhang, J. Wang, Y . Zhang, and L. Yang, “CycleResearcher: Improving automated research via automated review,” inInternational Conference on Learning Representations,

  15. [16]

    Available: https://arxiv.org/abs/2411.00816

    [Online]. Available: https://arxiv.org/abs/2411.00816

  16. [17]

    OpenLens AI: Fully autonomous research agent for health informatics,

    Y . Cheng and J. Suo, “OpenLens AI: Fully autonomous research agent for health informatics,”arXiv preprint arXiv:2509.14778, 2025. [Online]. Available: https://arxiv.org/abs/2509.14778

  17. [18]

    Accelerating scientific research with Gemini: Case studies and common techniques,

    D. P. Woodruff, V . Cohen-Addad, L. Jain, J. Mao, S. Zuo, M. Bateni, S. b. Branzei, M. P. Brenner, L. Chen, and Y . Feng, “Accelerating scientific research with Gemini: Case studies and common techniques,”arXiv preprint arXiv:2602.03837, 2026. [Online]. Available: https://arxiv.org/abs/2602.03837

  18. [19]

    EpidemIQs: Prompt-to-paper LLM agents for epidemic modeling and analysis,

    M. H. Samaei, F. D. Sahneh, L. W. Cohnstaedt, and C. M. Scoglio, “EpidemIQs: Prompt-to-paper LLM agents for epidemic modeling and analysis,”arXiv preprint arXiv:2510.00024, 2025. [Online]. Available: https://arxiv.org/abs/2510.00024

  19. [20]

    NORA: A harness-engineered autonomous research agent for end-to-end spatial data science,

    B. Zhou, X. Huang, H. Ning, Q. Wu, D. Li, and Z. Zhang, “NORA: A harness-engineered autonomous research agent for end-to-end spatial data science,”arXiv preprint arXiv:2605.02092, 2026. [Online]. Available: https://arxiv.org/abs/2605.02092 APPENDIXA GENERATEDMANUSCRIPTS: UPDATEDPIPELINE The following pages reproduce the complete RLEv4- generated manuscrip...