{"id":"0b0d4b9b-7be4-4b9c-8a9d-5a591c9b02d6","arxiv_id":"2607.03863","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SCION claims an agentic OS with Research Execution Plans and layered memory that beats autonomous research-agent baselines on reading, ideation, molecule design, and antibody screening.","lead":"SCION is a multi-agent “scientific operating system” that turns research goals into executable Research Execution Plans with memory, verification, and tool coordination. If it works as claimed, labs could automate more of the glue work between literature, models, simulation, and screening.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Reported gains may not be causally attributable to SCION’s OS design because backbone models, baseline adaptations, and LLM judging are uncontrolled.","rationale":"The paper’s contribution as an agentic OS framing (REP, hierarchical agents, layered memory, inverse/active-search formulations) is coherent and the empirical tables/figures show SCION ahead of listed baselines. The load-bearing issue is not that the architecture is incoherent, but that the strongest claim—that those OS mechanisms cause the gains—depends on comparisons that confound architecture with model choice, baseline adaptation, and LLM-as-judge scoring. The reader already identified this as the weakest assumption; the stress test finds the same soft spot and does not surface a deeper internal inconsistency in the math or architecture. A controlled fixed-backbone re-run would settle whether the OS claim survives. That keeps the verdict CONDITIONAL rather than ACCEPT or REJECT: promising systems results contingent on fairer evidence. No stronger objection (e.g., contradiction in the inverse-search formalization) is required to explain the current evidence gap.","tokens_in":21235,"tokens_out":604,"duration_ms":5227,"concrete_test":"Re-run the four benchmarks with one fixed backbone (e.g., nex-n1.1 or GPT-5.5) for SCION and every baseline, no task-specific rewiring beyond a shared tool API, and for Idea Maker add a second independent judge plus human novelty ratings on a 10-query subset; if SCION’s margins shrink below significance or win rates fall near chance under the fixed-backbone protocol, the OS-causal claim does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that SCION’s Meta-Harness / REP / multi-agent OS design causes outperformance on reading, ideation, molecule generation, and antibody screening (esp. decomposition, verification, refinement, memory reuse). That causal attribution is the least secure link. Section 7 mixes backbones (Kimi 2.5 for reading; nex-n1.1 for idea/molecule/antibody; Claude Opus 4.6 and GPT-5.5 for ARIS) and states that pipeline agents are adapted “to specified tasks,” so baselines are not fixed-backbone, fixed-protocol controls. Idea novelty (Table 6) is GPT-5.5 pairwise judging with 100% win rates against AutoResearchClaw and ScienceClaw and 98.33% vs EvoScientist—extreme margins that are consistent with judge preference or output-format bias rather than OS-driven novelty. Antibody screening (Fig. 7) uses synthetic intermediate-round feedback, reports a single F1 without error bars or multi-round budget curves, and no code/data are released. Without isolating architecture from model strength and evaluation artifacts, the OS-level causal claim remains underdetermined even if absolute scores favor SCION.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This paper presents SCION, an agentic scientific operating system that treats research workflows as first-class computational objects. A Science Agent Meta-Harness compiles high-level intent into Research Execution Plans (REPs) specifying stages, dependencies, verification checkpoints, tools, artifacts, and fallbacks, then coordinates hierarchical profile-driven agents, selective context, governed delegation, and layered L1–L3 epistemic memory. Discovery is formalized as target-conditioned inverse search (known targets) and as batch active search under finite budgets (hidden targets). Applications are sketched for materials analysis, multi-property molecule design, and antibody screening. Experiments on CMPhysBench scientific reading, idea novelty (LLM pairwise), multi-property molecule generation, and one-round antibody screening report gains over pipeline and agentic research baselines, attributed especially to decomposition, verification, refinement, and memory reuse.","tokens_in":21596,"tokens_out":1386,"duration_ms":14857,"significance":"If the architectural claims hold under controlled evaluation, SCION would be a meaningful systems contribution to AI4Science: it reframes the bottleneck as coordination, provenance, and recoverability rather than isolated model accuracy, and it offers a concrete OS-like substrate (REP, profile kernel, layered memory, governed delegation) that many current research agents lack. The inverse-search / batch-active-search framing is useful as an operational lens even though it is standard optimization language rather than a new theorem. Strengths include a coherent multi-agent design (Tables 1–4), explicit mapping of applications to formulations (Table 5), and multi-task empirical coverage. The work is significant primarily as systems architecture plus empirical systems evaluation; its lasting value depends on whether gains can be attributed to the OS design rather than backbone or evaluation artifacts.","major_comments":[{"comment":"Section 7 (Implementation details and all four benchmarks): The central causal claim—that SCION’s Meta-Harness / REP / multi-agent OS design drives outperformance in decomposition, verification, refinement, and memory reuse—is not isolated. Benchmarks use different backbones (Kimi 2.5 for reading; nex-n1.1 for idea/molecule/antibody; Claude Opus 4.6 and GPT-5.5 for ARIS), and pipeline agents are adapted “to specified tasks.” Without fixed-backbone, fixed-protocol controls and component ablations (REP off, memory off, verification off, single-agent vs hierarchical), absolute score advantages cannot be attributed to the organizational design. This is load-bearing for the abstract’s and conclusion’s OS-level claim.","section":null},{"comment":"§7.2 and Table 6: Idea novelty is evaluated solely by GPT-5.5 pairwise judgments, with SCION win rates of 100% vs AutoResearchClaw and ScienceClaw and 98.33% vs EvoScientist. Extreme margins under a single LLM judge are consistent with format, verbosity, or stylistic preference rather than OS-driven novelty. The paper needs either human expert ratings, multi-judge agreement, or an objective novelty proxy (e.g., literature-overlap / citation-grounded distinctness), plus disclosure of prompt templates and output-length controls, before novelty can support the architecture claim.","section":null},{"comment":"§5.2 and §7.4 / Figure 7: Batch active search is formulated as multi-round nonmyopic maximization of expected positives under budget T=Hb (Eqs. 17–18), yet the experiment evaluates only a single intermediate round with synthetic feedback, a fixed top-10% labeling rule, a 1:1 train/test split, and a single F1 (0.370) without error bars, multi-seed variance, multi-round budget curves, or comparison to standard active-search baselines (e.g., myopic top-p, nonmyopic search). This under-tests the hidden-target formulation that the paper presents as a core contribution.","section":null},{"comment":"§7.3 / Figure 6: Multi-property molecule success rates favor SCION on average, but the paper does not report candidate-set sizes, generation budgets, verifier definitions, or whether baselines had equal access to the same property predictors, validity filters, and iterative refine loops. Without matched tool surfaces and equal query budgets, higher success rates may reflect tool orchestration privileges rather than the REP/memory architecture per se. Equalize tools and report per-objective constraint satisfaction rates.","section":null}],"minor_comments":[{"comment":"Abstract and §1: Several compound words appear without spaces (e.g., “systemsremainfragmented,” “ScientificCollaborativeInnovation”). Clean typesetting throughout.","section":null},{"comment":"§5.1 Eqs. (12)–(13): f^{-1}_SA is defined as an informal runtime procedure then placed inside an arg min over procedure class F_inv; clarify that this is schematic notation, not a solved optimization, to avoid over-reading as a closed-form result.","section":null},{"comment":"Figures 5–7: Report numerical values in tables as well as bar charts; add error bars or multi-run statistics where stochastic generation or sampling is involved.","section":null},{"comment":"Related Work §2.2: Position more clearly against general LLM-agent OS work (e.g., AIOS) and against self-driving-lab literature beyond the brief future-work note in the conclusion.","section":null},{"comment":"Reproducibility: No code, prompts, REP schemas, or dataset splits are released. At minimum, provide REP schema examples, evaluation scripts, and antibody split definitions.","section":null},{"comment":"Table 1 vs text: “idea” and “reading” are listed as shared specialists while sci/exec are project-scoped; briefly justify why ideation and literature are not project-private when projects often need confidential literature and idea branches.","section":null}],"recommendation":"major_revision","confidential_remarks":"Fit is reasonable for a systems/AI4Science venue, but the manuscript currently reads more as an architecture-plus-benchmarks systems paper than as a theory paper; the inverse-search math should not be oversold. The skeptic’s concern about causal attribution is well-founded and should drive revision priority. If the authors can add fixed-backbone ablations and stronger human or multi-judge idea evaluation, the contribution becomes much more credible; without that, risk of overclaiming OS causality remains high."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a real systems paper, not vapor. SCION packages multi-agent research automation as an organizational OS around Research Execution Plans, a Science Agent meta-harness, profile-driven specialists, selective context, and layered memory. That packaging is the contribution. The inverse-search / batch-active-search math is standard language, not a new theorem, but it is a clean way to state what the runtime is trying to do.\n\nWhat it does well: the architecture section is careful and usable. Treating the research plan as a first-class object with stages, dependencies, verifiers, artifacts, and fallbacks is the right systems move. Related work is honest about ChemCrow, Coscientist, AI-Scientist, InternAgent, AIOS, and the rest. On the empirical side, multi-property molecule generation is the most persuasive result—success rates rise where simultaneous validity, filtering, and multi-objective ranking matter, which matches the verification/refinement story. Reading gains on CMPhysBench are real but small. Antibody F1 is better than the named baselines in the reported one-round setup.\n\nSoft spots, in proportion: the stress-test concern lands. Backbones differ (Kimi 2.5, nex-n1.1, Claude/GPT for ARIS), pipeline agents are adapted “to specified tasks,” idea novelty is GPT-5.5 pairwise judging with near-100% win rates against several systems, antibody uses synthetic intermediate-round feedback without error bars or multi-round curves, and no code or data are released. So absolute scores favor SCION; attributing that to the OS design rather than model strength, adaptation, or judge format is not yet secured. That is a medium flaw for a systems claim, not a reason to dismiss the design.\n\nWho it is for: people building research agents, lab automation stacks, or AI4Science infrastructure. If you care about long-horizon recoverability and plan-level memory, read Sections 4–5 and the molecule experiment. I would send it to peer review. Engage if that is your lane; treat the causal OS claim as provisional until controlled ablations and artifacts exist.","headline":"Useful OS packaging for agentic science work; the architecture is coherent, but the causal claim that SCION’s design drives the wins is still underdetermined.","tokens_in":22324,"tokens_out":538,"would_cite":false,"duration_ms":12100,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"SCION turns scientific research into an executable operating system: a Science Agent Meta-Harness compiles intent into Research Execution Plans and coordinates agents, tools, and memory, beating existing research-agent baselines on multi-st","keywords":["agentic AI","scientific discovery","Research Execution Plan","multi-agent systems","AI4Science","inverse design","batch active search","epistemic memory"],"falsifier":"Fix one shared model backbone and identical task adapters, then ablate REP compilation, layered memory, and governed verification on the multi-property molecule-generation and one-round antibody-screening suites; if success rate and F1 drop to match strong baselines, the organizational claim holds; if performance stays high, the OS design is not what carries the gains.","tokens_in":22107,"feed_emoji":"🔬","tokens_out":956,"duration_ms":16129,"temperature":0.7,"pith_summary":"Most AI-for-science systems still leave humans as the glue between problem formulation, literature, models, simulation, validation, and reuse. This paper introduces SCION as an agentic scientific operating system—an organizational nexus—that makes the research workflow itself a first-class computational object. A Science Agent acting as a Meta-Harness compiles high-level intent into a Research Execution Plan with stages, dependencies, verification checkpoints, tools, expected artifacts, and fallbacks, then runs hierarchical multi-agent execution with selective context, governed delegation, and layered epistemic memory. Discovery is cast as target-conditioned inverse search and, when the target is hidden, as batch active search under a finite experimental budget. On scientific reading, idea generation, multi-property molecule generation, and antibody screening, SCION is reported to outperform pipeline and agentic research baselines especially where decomposition, verification, refinement, and memory reuse matter.","feed_headline":"AI as a lab OS beats research-agent baselines","feed_subtitle":"Plans, agents, verification, and memory lift reading, ideas, molecules, and antibody screens.","key_machinery":"The Research Execution Plan (REP): a structured, machine-operable plan that compiles high-level scientific intent into staged objectives, dependencies, verification checkpoints, tool requirements, expected artifacts, and fallback conditions. The Science Agent Meta-Harness uses the REP to coordinate agents, tools, and layered memory as an approximate inverse-search procedure over candidates and experimental trajectories.","core_discovery":"The paper claims that scientific discovery improves when AI is no longer a set of isolated tools but a coordinated operational layer: compiling scientific intent into Research Execution Plans, executing them through a hierarchical multi-agent runtime with verification and recovery, and storing trajectories as reusable epistemic memory yields more auditable, recoverable work and higher performance than existing autonomous research agents on multi-stage scientific tasks.","pith_inferences":["If REP-style plans become common, labs will need shared schemas for verification checkpoints and provenance before multi-group collaboration scales cleanly.","The same batch-active-search harness should transfer to other sparse-hit discovery settings (for example catalyst or assay screening) without inventing a new architecture.","The largest reported lifts on multi-constraint molecule tasks suggest one-shot generators remain weak where validity filters and trade-offs must be enforced online.","A stripped planner-plus-memory layer on a fixed backbone would be a strong control experiment to isolate how much of the gain is truly organizational."],"forward_implications":["Human scientists can move from manual tool dispatch toward strategy, value constraints, and high-level judgment while the system handles coordination and recovery.","Materials analysis, multi-property molecule design, and antibody screening can be run as governed inverse-search or budgeted batch-active-search procedures with auditable branches.","Failed attempts, intermediate artifacts, and decision traces become reusable project memory instead of discarded session noise.","Evaluation of AI for science expands from single-model accuracy to system-level recoverability, auditability, and knowledge reuse.","Coupling this operating-system model with automated experimentation and wet labs can form a cyber-physical infrastructure for science."],"fun_headline_variants":["SCION turns science into executable plans and beats research agents","Agentic OS for science: plans, agents, memory outperform baselines","From tools to lab OS: SCION lifts multi-stage discovery tasks","Research Execution Plans make AI discovery auditable and stronger","Hierarchical agents plus memory beat autonomous research baselines"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The reported gains come from SCION’s organizational design rather than from stronger or differently prompted model backbones, unequal baseline adaptations, or the preferences of the LLM judge used for idea novelty.","fun_headline_variants_meta":{"raw":{"variants":["SCION turns science into executable plans and beats research agents","Agentic OS for science: plans, agents, memory outperform baselines","From tools to lab OS: SCION lifts multi-stage discovery tasks","Research Execution Plans make AI discovery auditable and stronger","Hierarchical agents plus memory beat autonomous research baselines"]},"model":"grok-4.5","effort":"low","cost_usd":0.004784,"raw_usage":{"total_tokens":1404,"prompt_tokens":814,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":47840000,"prompt_tokens_details":{"text_tokens":814,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":525,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":814,"tokens_out":65,"duration_ms":4307,"temperature":1.0,"reasoning_tokens":525,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T23:25:08.418477+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Fix one shared model backbone and identical task adapters, then ablate REP compilation, layered memory, and governed verification on the multi-property molecule-generation and one-round antibody-screening suites; if success rate and F1 drop to match strong baselines, the organizational claim holds; if performance stays high, the OS design is not what carries the gains.","supporting_citations":[],"review_version":1}