{"id":"80aaf56a-8d3a-4ca1-bd3c-984c3163f555","arxiv_id":"2607.27556","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Agentic bioinformatics systems mostly demonstrate planning and tool execution but rarely prospective empirical validation, so the paper argues evaluation should center on inspectable workflow trajectories (FEV) rather than final-answer correctness.","lead":"A review of 128 papers on AI agents for biology proposes judging them by the workflow they run — what they do, what evidence backs them, and how the result was tested — rather than by the final answer alone. It finds most agents can plan and execute tools, but only about 6 percent have been tested in real experiments and almost none close the loop with lab feedback.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"V2/V3 coding may overstate replayability: entries like AI-HOPE and HEAL-KGGen are assigned V2 on 'public code/instructions' without the Table S5-required intermediate artifacts/execution traces, and no actual replay or inter-rater reliability check is reported.","rationale":"The reader's weakest assumption is correct and is the load-bearing point. There is a real inconsistency between the stated V2 gate and several exemplar entries, and no reliability check is reported. However, this is a methodological weakness in the empirical map, not a flaw in the paper's central normative argument. The core thesis—evaluate workflow correctness rather than final-answer accuracy alone—is well-motivated and does not depend on exact counts. Even a substantially reduced V2/V3 count would leave the qualitative lopsidedness (only 7 V4, closed-loop refinement in 1 system) intact. The paper also discloses the sample-based, reporting-dependent nature of its counts (Section S1), which mitigates selection concerns. Duplicate reference [49]=[50] is a minor mechanical issue. Therefore I keep the reader's CONDITIONAL verdict (no verdict change), with the concrete test above as the requested robustness check. I agree with the reader's identification of the weakest assumption.","tokens_in":46281,"tokens_out":7899,"duration_ms":80048,"concrete_test":"Independently sample 30 of the 94 entries coded V2/V3 (deliberately including AI-HOPE, HEAL-KGGen, SwiftDossier). Two coders re-apply the Table S5 criteria using only the cited primary sources; compute Cohen's kappa. Then attempt actual replay of each still-V2/V3 workflow from its public artifacts (clone code, install dependencies, run with the cited inputs/parameters, compare generated intermediate artifacts and traces). If the strict recoding downgrades at least 20% of sampled systems below V2, or if at least half of attempted replays fail, the reported V-stage distribution and the abstract's replayability claim require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's quantitative synthesis depends on the V-stage assignments. Table S5 defines V2 as requiring identifiable inputs, parameters, dependencies, intermediate artifacts, AND execution traces. Several table entries apply a looser bar. AI-HOPE (Table S8) is coded V2 [S] on the strength of 'executable analyses on identifiable retrospective datasets,' with no listed scripts, parameter records, logs, or traces. HEAL-KGGen (Table S9) is coded V2 [B] because 'public code, requirements, test data, graph files, and instructions support replay'; instructions and graph files are not execution traces, and no intermediate artifacts are listed. SwiftDossier (Table S23) is coded V2 [H,R] from 'executable retrieval and analysis artifacts' without the required intermediate artifacts/traces. Because V3 subsumes V2, the 69 'V3' entries inherit this slack. Section S1 reports no inter-rater reliability or calibration exercise, and the manuscript gives no evidence that any workflow was actually replayed; V2 is inferred from the publication description rather than demonstrated. The paper's own caveat that unreported capabilities were not coded does not fix this: the flagged entries were coded leniently relative to the stated gate. A stricter re-application of Table S5 would shift the V-stage distribution, lower the 86% replayable-computation figure, and undermine the 'most systems are classified as V3' summary. The normative core (assess workflow correctness) is not damaged by this; the empirical map is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Function–Evidence–Validation (FEV) framework for evaluating agentic bioinformatics systems. Function records demonstrable workflow operations (F1–F6), Evidence records traceable sources supporting actions and claims (E1–E6), and Validation records cumulative assurance stages (V0–V4) with orthogonal qualifiers. The authors apply FEV to 109 system entries and 28 benchmark resources, representing 128 unique publications, and present a cross-domain synthesis showing that planning and tool-mediated execution are common while replayability, external validation, and prospective empirical testing are less well established. They conclude that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone.","tokens_in":46483,"tokens_out":7772,"duration_ms":86853,"significance":"The paper makes a timely and useful conceptual contribution. The FEV framework separates operational capability, evidentiary support, and validation assurance in a way that is more explicit than most existing reviews, and the supplementary tables provide an unusually detailed audit trail of 109 systems. The distinction between using empirical data as evidence and prospectively testing an agent-generated output is particularly valuable, as is the cumulative V0–V4 gate. The accounting is internally consistent (109 + 28 − 9 = 128 unique publications), and the authors are transparent about the review's scope and limitations. The main risk is that the quantitative V-stage synthesis is not calibrated: several V2/V3 assignments appear more lenient than the stated criteria, and no inter-rater reliability or replay check is reported. This weakens the specific numeric claims but does not invalidate the normative core, which would only be strengthened if replayability is even rarer than reported.","major_comments":[{"comment":"The V2 gate is defined in Table S5 as requiring identifiable inputs, parameters, dependencies, intermediate artifacts, AND execution traces. Several assignments use a looser bar. AI-HOPE (Table S8) is coded V2 [S] on 'executable analyses on identifiable retrospective datasets,' with no listed scripts, parameter records, logs, or traces. HEAL-KGGen (Table S9) is coded V2 [B] because 'public code, requirements, test data, graph files, and instructions support replay'; instructions and graph files are not execution traces. SwiftDossier (Table S23) is coded V2 [H,R] from 'executable retrieval and analysis artifacts' without the required intermediate artifacts/traces. Since V3 subsumes V2, the 69 V3 entries inherit this slack. Section S1 reports no inter-rater reliability or calibration exercise and no actual replay check. This undermines quantitative statements such as 'most systems are clas","section":"§S2 Tables S8, S9, S23; Table S5; §S1"},{"comment":"The eligibility boundary is deliberately broad, including 'agent-adjacent' systems, and §S1 states that aggregate analyses include both unless otherwise stated. However, the main quantitative synthesis does not report the full-agentic subset separately. Entries such as ChatNT (Table S7), VibeGen (Table S20), and ORI (Table S21) are coded as agent-adjacent predictive or model–laboratory loops, yet they contribute to the same FEV prevalence and V-stage distribution as full multi-agent workflow systems. The claim that 'planning and tool-mediated execution have advanced' across agentic bioinformatics is therefore hard to interpret. Please report the full-agentic-only distribution or provide a sensitivity analysis showing that the qualitative conclusions are unchanged when agent-adjacent entries are excluded.","section":"§2 and §S1; Figures S3–S5"}],"minor_comments":[{"comment":"References 49 and 50 are identical (Huang et al., 'Autonomous biomedical research with an artificial intelligence agent'). Please consolidate or distinguish them. References 117 and 118 also appear to be two versions of BioMaster and should be cross-referenced explicitly.","section":"References"},{"comment":"The zero count for V0 is partly an artifact of the eligibility filter that excludes purely conversational systems. Add a note clarifying that V0 is retained on the complete scale but that no system meeting the inclusion criteria was assigned to it.","section":"Figure S3b"},{"comment":"The term 'workflow correctness' is used in the abstract and conclusion but never given a compact definition in the body. A brief formal definition early in Section 2 would help readers understand the exact relationship between FEV and workflow correctness.","section":"Abstract and §2"},{"comment":"The 'observed gap' score (1 − share) measures absence of reporting, not absence of capability. Since the paper codes only reported capabilities, consider relabeling this as a 'reporting gap' or adding an explicit sentence that unreported capabilities were treated as not demonstrated.","section":"Figure S5b"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and potentially influential review. The FEV framework is a genuine contribution, and the supplementary tables are unusually transparent. My main concern is the uncalibrated V2/V3 coding, which affects the quantitative synthesis but not the normative thesis. I would support acceptance after the authors either re-code the V-stage assignments with a stricter, reproducible protocol or substantially soften the numeric claims. The paper fits the journal's scope, and the duplicate reference issues are easily fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about arXiv:2607.27556. One: the FEV framework — separating Function, Evidence, and Validation for agentic bioinformatics workflows — is a genuinely useful contribution and the paper is worth reading. Two: the empirical V-stage coding is sloppier than the stated criteria, so the headline counts (109 systems, 86% at V2+, 69 at V3, 7 at V4) should be read as indicative, not exact.\n\nWhat's actually new: treating the inspectable workflow trajectory as the unit of analysis, not architecture or final answer. The paper makes the case that benchmark accuracy and tool-call fluency don't establish scientific credibility, and that's well argued. The V0–V4 validation stages with qualifiers (B, H, S, R, X, P, C) give the field a common vocabulary. The cross-domain synthesis is a solid piece of work; the accounting checks out (109 + 28 − 9 = 128, and the V-stage sum is correct). The finding that only 7 of 109 systems reach prospective empirical evaluation, and only 1 has closed-loop refinement, is a useful corrective to the hype.\n\nThe soft spots are real but not fatal. The stress-test is right: Table S5 requires V2 to include identifiable inputs, parameters, dependencies, intermediate artifacts, and execution traces. Yet several table entries — AI-HOPE, HEAL-KGGen, SwiftDossier — are assigned V2 on \"public code\" or \"executable analyses\" without the required artifacts and traces. Since V3 subsumes V2, the 69 V3 count inherits the same slack. No inter-rater reliability check or actual replay verification is reported. So the \"replayability gap\" is overstated; with a strict reading, more systems drop to V1. The normative argument survives, but the empirical map needs tightening. Also, the abstract bundles the weak replayability gap with the strong prospective-testing gap; those should be separated. Minor: reference [49] is duplicated as [50].\n\nWho benefits: anyone building or evaluating bioinformatics agents, and anyone writing benchmarks. The framework is a good scaffold even if the assignments need calibration.\n\nRecommendation: send it to peer review. A serious referee should ask for a coding log, a stricter and consistent V2 gate, and softened claims about the replayability gap. The core idea is sound and the resource is timely.","headline":"Useful evaluation scaffold for agentic bioinformatics; the empirical V-stage coding is looser than the stated gate, so treat the counts as indicative.","tokens_in":47140,"tokens_out":2840,"would_cite":true,"duration_ms":28862,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI biology agents should be judged by workflow correctness, not final answers alone.","keywords":["agentic bioinformatics","workflow correctness","Function–Evidence–Validation","validation stages","evidence grounding","replayability","large language model agents","scientific accountability"],"falsifier":"Re-code the 109 mapped systems from the paper's own tables using only the explicit V2 minimum (inputs, parameters, dependencies, intermediate artifacts, and execution traces all present). If a stricter coder assigns substantially fewer than 94 systems to V2—for instance, because 'public code and instructions' is treated as insufficient without traceable execution artifacts—the claimed replayability gap and the V-stage distribution would shift together.","tokens_in":45965,"feed_emoji":"🧬","tokens_out":3557,"duration_ms":35362,"temperature":0.7,"pith_summary":"The paper argues that large language model agents in bioinformatics are being evaluated by the wrong yardstick. Fluent outputs, successful tool calls, and benchmark scores say little about whether the process that produced them was scientifically defensible. The authors propose that the inspectable workflow trajectory—the sequence of objectives, data, decisions, tools, artifacts, and evidence—should be the primary unit of analysis, and they introduce a Function–Evidence–Validation (FEV) framework to measure it. Applying FEV to 109 systems, they find that planning and tool use have advanced far ahead of replayability, provenance, and prospective empirical testing: 94 of 109 systems meet the replayability gate, but only 7 have any prospective experimental validation. The conclusion is that the field should shift from final-answer correctness to workflow correctness.","feed_headline":"Judge AI bioinformatics by workflow, not just final answers","feed_subtitle":"A new framework rates 109 systems on function, evidence, and validation; only 7 earn prospective empirical testing.","key_machinery":"The Function–Evidence–Validation (FEV) framework. Function records what the system demonstrably does (planning, coordination, tool execution, state and trace maintenance, repair, verification). Evidence records traceable sources (literature, knowledge bases, measurements, software outputs, model outputs, experimental observations). Validation is a five-stage cumulative ladder from illustrative output (V0) to demonstrated execution (V1), replayable computation (V2), scientifically evaluated computation (V3), and prospective empirical evaluation (V4). The ladder is the load-bearing instrument: it separates 'it ran' from 'it can be replayed' from 'it was scientifically tested.'","core_discovery":"The central claim is normative: agentic bioinformatics should be evaluated through workflow correctness rather than final-answer correctness alone. The paper operationalizes this as three non-interchangeable properties—demonstrated workflow operations (Function F1–F6), traceable support for actions and claims (Evidence E1–E6), and use-case-specific cumulative assurance (Validation V0–V4)—and maps 109 systems plus 28 benchmarks across six biological domains. The empirical finding is lopsided progress: 94 of 109 systems reach the V2 replayability gate, but only 7 reach V4 prospective empirical testing, and closed-loop empirical refinement appears in a single mapped system. If the paper is righ","pith_inferences":["If FEV became a reporting norm, the field's progress measures would shift from 'can it answer' to 'can it be audited'—a change that would likely re-rank many systems and reduce the prestige of broad-but-untested agents.","The V-ladder could be extended with a V5 for closed-loop empirical refinement at scale, since the paper counts only one such system; the distinction between one-shot prospective testing (P) and feedback-driven cycles (C) is likely to become central as wet-lab integration grows.","A testable extension would be an inter-rater reliability study of FEV coding: if independent coders disagree widely on V-stage assignments, the framework needs tighter operational definitions before it can serve as a community standard."],"forward_implications":["Benchmarks in agentic bioinformatics should grade trajectories—tool calls, parameters, artifacts, failure recovery—alongside final answers, rather than treating endpoint accuracy as the whole score.","Published system claims should carry an explicit V-stage and qualifiers (benchmark, expert, statistics, robustness, external, prospective, closed-loop), so a reader can see what assurance a given use case actually has.","The replayability gap (94 of 109 at V2, 7 at V4) implies that most current systems are inspectable but not empirically tested; funding and evaluation efforts should move toward prospective, claim-aligned experiments.","A minimum reporting standard for agentic bioinformatics papers would follow from FEV: document workflow scope, models, tools, parameters, environments, artifacts, failures, approval points, and the validation stage."],"fun_headline_variants":["Only 7 of 109 AI bio agents survive full scrutiny","AI bio: workflow is the new truth, not the answer","Judge AI bio by how it works, not what it says","FEV framework: 109 AI bio tools, 7 proven","In AI bio, process beats product: the FEV test"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline numbers (94 of 109 at V2, 7 at V4) rest on the authors applying their own stated V2 criteria—identifiable inputs, parameters, dependencies, intermediate artifacts, and execution traces—consistently across 109 heterogeneous papers, and that coding has not been checked by independent raters.","fun_headline_variants_meta":{"raw":{"variants":["Only 7 of 109 AI bio agents survive full scrutiny","AI bio: workflow is the new truth, not the answer","Judge AI bio by how it works, not what it says","FEV framework: 109 AI bio tools, 7 proven","In AI bio, process beats product: the FEV test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000811,"raw_usage":{"total_tokens":3398,"prompt_tokens":752,"completion_tokens":2646,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":2559}},"tokens_in":496,"tokens_out":2646,"duration_ms":20258,"temperature":1.0,"reasoning_tokens":2559,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:43:18.534070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the 109 mapped systems from the paper's own tables using only the explicit V2 minimum (inputs, parameters, dependencies, intermediate artifacts, and execution traces all present). If a stricter coder assigns substantially fewer than 94 systems to V2—for instance, because 'public code and instructions' is treated as insufficient without traceable execution artifacts—the claimed replayability gap and the V-stage distribution would shift together.","supporting_citations":[],"review_version":1}