{"id":"fd888dc8-ed36-4ecb-b551-47a72b8b16d7","arxiv_id":"2607.05696","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Autonomous controllers cannot remove decision risk from target components absent from their pre-action physical records; closing that gap requires expanded access, cost, or restricted deployment.","lead":"This paper frames autonomous science as limited by which physical measurements a platform can expose before it acts, not just by algorithms. It defines a Bayes-optimal risk frontier over constrained records and an audit that flags missing target-relevant variables before deployment.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified to the load-bearing projection claim; the reader's caveat is real but non-undermining.","rationale":"The reader's strongest claim is exactly Theorem 1: the Bayes residual under the autonomous pre-action record is a hard floor for any controller using only that record. That is standard orthogonal projection / conditional-variance geometry; the paper's contribution is packaging it as a constrained measurement-access frontier and pre-deployment audit (PADM), not inventing a new inequality. The reader's weakest_assumption correctly notes that real autonomous labs face misspecification and learning error, but the manuscript already treats those as additive above the access bound and does not claim the frontier is always the binding operational limit. Therefore the concern does not move the verdict: CONDITIONAL remains appropriate because of non-public code at submission, oracle/epsilon dependence, and the need for clearer real-data estimation protocols, not because the projection identity is insecure. Agreement with the reader is full on both the claim and the caveat; no stronger load-bearing attack on the math is warranted. The concrete test simply re-checks the cleanest analytic equality case the paper itself uses to illustrate the floor.","tokens_in":13129,"tokens_out":647,"duration_ms":5484,"concrete_test":"Independently re-derive the residual risks in the hidden-regime benchmark of Fig. 4 / Supp. Sec. 4: with T=beta z x, Y_A=x, cue reliability p, confirm R_A=beta^2 E[x^2], R_{A+H}=4p(1-p) beta^2 E[x^2], R_oracle=0, and that Monte Carlo at the stated p grid matches within sampling error. If those equalities fail under the paper's own generative model, the projection floor is misstated; if they hold, the strongest claim stands and only the diagnostic-scope caveat remains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 1 / Corollary 1) is the classical L2 projection identity: under quadratic loss, no square-integrable a(Y_A) can beat residual variance of T given the pre-action record. That identity is correct as stated, and the paper is explicit that finite-sample error, misspecification, learned representations, approximate inference and optimizer failure are separate computational limits that can only raise realized risk above the access floor (Discussion; Methods). The reader's weakest_assumption therefore does not falsify the theorem or the no-free-autonomy reading; it only questions whether the known-model oracle gap is the binding pre-deployment diagnostic when those other errors dominate. That is a scope and adoption caveat, not an internal inconsistency in the strongest claim. The synthetic checks (hidden-regime closed form, Gaussian precision update, monitored-feedback residual) are consistent with the identity. The remaining practical soft spot is estimation of Gamma_or and channel gains from pilot data under misspecification, which the paper already flags rather than conceals.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript formulates physically accessible decision-making (PADM) and a measurement-access risk frontier R*_auto(\\Lambda): the Bayes-optimal target risk minimized over records realizable under cost, bandwidth, latency, disturbance, memory and actuation constraints. Its central claim is a no-free-autonomy limit: under quadratic loss, any square-integrable action based only on the pre-action autonomous record Y_A cannot beat residual variance of the target T given Y_A (Theorem 1), so an optimal controller cannot remove target components absent from its record (Corollary 1). Closing the gap requires expanded access, auditing, tolerated disturbance, slower operation or restricted deployment. The claim is illustrated by monitored feedback with a hidden switching force, a chemistry-aware candidate-ranking audit with a 1000-target stress panel, Gaussian sensing, hidden-regime decisions and cost-aware/thermodynamic channel selection, plus an operational audit algorithm.","tokens_in":13469,"tokens_out":1100,"duration_ms":9285,"significance":"If the framing holds as a pre-deployment diagnostic, it supplies a clean architecture-level language for when autonomous scientific platforms are measurement-insufficient rather than merely under-optimized. The paper is explicit that Theorem 1 is the classical L2 projection identity and that the contribution is the frontier over constrained measurement architectures, residual oracle gaps and the audit workflow. Strengths include closed-form residual-risk formulas and equality cases for the Gaussian and hidden-regime benchmarks, a deterministic 1000-target chemistry stress panel with explicit recovery counts, monitored-feedback residual risk checked against the projection identity, and a stated retrospective CSV audit protocol with a public CAMEO/NIST shadow-mode check. These make the access floor falsifiable in the known-model setting and useful as a preflight check alongside experimental design, POMDPs and RL rather than as a replacement for them.","major_comments":[{"comment":"Discussion and Methods: the paper correctly states that finite-sample error, misspecification, learned representations, approximate inference and optimizer failure can only raise realized risk above the access floor, but the operational claim that PADM identifies residual oracle gaps before deployment still depends on estimating \\Gamma_or and channel gains G_j / A_j from pilot data. The manuscript should state more explicitly under what pilot-data conditions those estimators remain informative (e.g., blocked splits, negative controls, distribution shift), and what decision the audit should return when estimation noise or misspecification is comparable to the reported gap, so that the diagnostic is not over-read as binding whenever computational errors dominate.","section":null},{"comment":"Section 1.1, Eq. (1) and the measurement-sufficiency condition: R*_auto(\\Lambda) is defined as an infimum over the feasible autonomous record set Y_auto_acc(\\Lambda), yet the main demonstrations fix particular architectures (displacement-only vs cue; descriptor-only vs chemistry audit) rather than characterizing or approximating that set under concrete \\Lambda. A short constructive statement of how Y_auto_acc is delimited for at least one platform class (e.g., monitored feedback with bandwidth/latency, or the chemistry audit with pre-action timing) would make the frontier operational rather than primarily conceptual.","section":null}],"minor_comments":[{"comment":"Figure 2c: the sufficiency contour \\Delta R=0.10 is illustrative; state in the caption or Methods that it is a chosen threshold, not a universal criterion, and how it maps to the task-specific \\epsilon in the sufficiency definition.","section":null},{"comment":"Figure 3 and the chemistry panel: emphasize earlier in the main text (not only in the caption and Discussion) that this is a fixed-seed adversarial ranking stress test, not experimental chemical validation or a competitive generator benchmark.","section":null},{"comment":"Algorithm 1: the recovered-oracle-gap fraction A_j is undefined when R(Y_A)=R(Y_or); the skip rule is stated, but a one-line note that A_j is then omitted (or set to 1 by convention) would avoid implementation ambiguity.","section":null},{"comment":"Notation: Y_auto_acc(\\Lambda), P_adm(Y), \\Gamma_or and R_dyn(Y) are introduced densely in Section 1.1–1.2; a short symbol table or consistent first-use expansion would help readers outside mathematical physics.","section":null},{"comment":"Code availability: the manuscript states code will be public upon publication and is available to editors/reviewers on request; for reproducibility of the 1000-target panel and unit tests, a stable archive identifier at acceptance would strengthen the claim.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Fit is interdisciplinary (math-ph framing applied to autonomous science). The novelty is architectural rather than a new inequality; that is disclosed, but some venues may still ask whether the contribution is primarily conceptual packaging. The central math is sound; I would not reject on novelty grounds if the journal values pre-deployment diagnostics for autonomous labs. No citation-pattern concern stood out."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The load-bearing claim is correct and not oversold: under quadratic loss, no controller using only the pre-action record YA can beat residual variance of T given YA (Theorem 1 / Corollary 1). That is the classical L2 projection identity, and the paper is explicit that it is classical. What is new is the packaging: a constrained feasible-record frontier R*_auto(Λ), the no-free-autonomy slogan, and a target-specific pre-deployment audit (PADM / Algorithm 1) that treats measurement architecture as a first-class risk object for autonomous science.\n\nThe paper does this cleanly. The monitored-feedback example, Gaussian precision update, and hidden-regime closed forms recover the expected residual-risk formulas and equality cases. The chemistry stress panel is deterministic and adversarial by construction (descriptor-only selects the artifact 1000/1000; audit recovers the target), so it checks the audit logic rather than claiming discovery. Citations cover decision theory, Blackwell, experimental design, POMDPs, and quantum filtering without obvious gaps. The Discussion correctly separates the access floor from finite-sample, misspecification, and optimizer errors that can only raise realized risk.\n\nSoft spots are real but secondary. Code is not public at submission (available on request). Oracle proxy and sufficiency tolerance ε are task-dependent, so the diagnostic is only as good as those choices. The known-model Bayes residual may not be the binding limit when learning and model error dominate; that is a scope caveat the paper already flags, not an internal contradiction. Free parameters (ε, cue reliability, sensing costs) are stated and used for illustration.\n\nThis is for people building or auditing self-driving labs and closed-loop controllers who need a preflight check on whether the record stream can contain the variables the target requires. It is not a new theorem in decision theory. I would send it to peer review: formally grounded, reproducible analytic intent, and useful framing even if referees push for code release and clearer real-data estimation protocols. Worth engaging if you work on autonomous experimental platforms.","headline":"Solid architectural packaging of a classical Bayes projection fact into a pre-deployment measurement-access audit for autonomous labs; math is correct, novelty is framing not foundations.","tokens_in":14019,"tokens_out":519,"would_cite":true,"duration_ms":4725,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Autonomous science cannot collapse decision uncertainty by computation alone; residual risk is set by which physical records a platform exposes before it acts.","keywords":["physically accessible decision-making","measurement-access risk frontier","no-free-autonomy","autonomous science","Bayes risk floor","monitored feedback","sensor audit","hidden regimes"],"falsifier":"On a fixed real autonomous-lab target, add a candidate audit channel that should change the target projection; if held-out decision risk does not fall relative to the autonomous record alone while computational limits are controlled, the access floor is not the binding limit the audit claims to identify.","tokens_in":14027,"feed_emoji":"🔬","tokens_out":981,"duration_ms":19559,"temperature":0.7,"pith_summary":"This paper argues that scaling autonomous science is limited not only by algorithms, compute, or data volume, but by which physical records a platform can actually generate before it chooses an action. The authors formulate physically accessible decision-making (PADM) and a measurement-access risk frontier: the Bayes-optimal target risk minimized over records realizable under cost, bandwidth, latency, disturbance, memory, and actuation constraints. From this they derive a no-free-autonomy limit: an optimal controller cannot remove target components that never enter its record, so residual risk is an access floor rather than an optimizer failure. Closing that gap requires expanded sensing, auditing, tolerated disturbance, slower or staged operation, or restricted deployment. Concrete checks include monitored feedback with a hidden switching force, Gaussian and hidden-regime benchmarks, cost-aware and thermodynamic channel selection, and a chemistry-aware ranking audit on a 1000-target stress panel.","feed_headline":"Automation can't erase what the sensors never see","feed_subtitle":"A measurement-access risk frontier shows residual decision error is a physical floor, not a software bug.","key_machinery":"The measurement-access risk frontier R*_auto(Λ), carried by the autonomous-record risk floor (Theorem 1): for a scalar target T and pre-action record Y_A, under quadratic loss any square-integrable action a(Y_A) has risk at least E[(T − E[T|Y_A])²], with equality at the Bayes action. The projection identity converts residual risk into a diagnostic of missing physical access, and complementary-channel value equals how much the added record changes that target projection.","core_discovery":"The central claim is that autonomous scientific control is bounded by a measurement-access risk frontier: the best Bayes risk attainable over physically realizable pre-action records under platform constraints. Under quadratic loss, no controller using only the autonomous record can beat the residual variance left after projecting the target onto that record. That residual is therefore a physical access floor, not a computational defect, and the gap to an oracle closes only by expanding access, adding audit channels, slowing the loop, or restricting the operating domain.","pith_inferences":["Self-driving laboratories may need formal measurement-sufficiency checks before full closed-loop deployment, analogous to safety cases for acting under partial observability.","Many unexpected failures in robotic chemistry and materials loops could reclassify as missing channels rather than weak acquisition functions or under-trained policies.","Quantum feedback already distinguishes monitored channels from unmonitored modes; the same target-risk recovery criterion could rank channels under back-action cost.","If access gaps routinely dominate learning error in pilots, investment should shift from larger models toward instrument architecture and pre-action audit sensors."],"forward_implications":["Pre-deployment audits should score candidate sensors by target-specific risk recovery, not by generic information content.","A platform that acts before exposing a target-relevant mode retains irreducible residual risk no matter how strong the learning algorithm.","Closing an automation gap requires expanded sensing, audit channels, slower staged operation, or a restricted autonomous domain.","Redundant records leave the risk floor unchanged; only complementary access that moves the Bayes projection of the target reduces risk.","Cost, bandwidth, latency, and back-action enter as constraints on the feasible record set, so instrument and thermodynamic limits become part of the decision bound."],"fun_headline_variants":["No free autonomy: missing records set a hard risk floor","Controllers can't erase target parts sensors never capture","Decision risk floors on physical access not pure compute","Bayes risk can't beat residual variance from the record","Measurement access not algorithms bound autonomous control"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The paper treats the known-model Bayes residual under a chosen target and an idealized oracle as the right pre-deployment diagnostic, even though real systems also face model error, finite samples, and approximate inference.","fun_headline_variants_meta":{"raw":{"variants":["No free autonomy: missing records set a hard risk floor","Controllers can't erase target parts sensors never capture","Decision risk floors on physical access not pure compute","Bayes risk can't beat residual variance from the record","Measurement access not algorithms bound autonomous control"]},"model":"grok-4.5","effort":"low","cost_usd":0.006114,"raw_usage":{"total_tokens":1581,"prompt_tokens":747,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":61140000,"prompt_tokens_details":{"text_tokens":747,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":762,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":747,"tokens_out":72,"duration_ms":7847,"temperature":1.0,"reasoning_tokens":762,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T03:34:58.725866+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a fixed real autonomous-lab target, add a candidate audit channel that should change the target projection; if held-out decision risk does not fall relative to the autonomous record alone while computational limits are controlled, the access floor is not the binding limit the audit claims to identify.","supporting_citations":[],"review_version":1}