{"id":"908b06b9-dd86-4945-91da-9a3f04710dda","arxiv_id":"2607.25045","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An MNE-grounded EEG agent separates LLM semantic routing from deterministic commitment, participant-disjoint confirmation, and evidence-bound release, outperforming a matched lexical router and blocking adaptive-search hazards in sealed tests.","lead":"CogEEGAgent is an LLM agent for cognitive EEG that lets the model choose a registered analysis from natural language while deterministic code owns contracts, held-out confirmation, and release. It matters as a concrete pattern for fail-closed scientific agents where fluent reports are not enough.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The \"fail-closed\" guarantee is process-closed, not intent-closed: the only empirical test of semantic misfit (3 underspecified requests) failed 3/3 and was contained solely by request ordering, which the authors concede is not order-invariant.","rationale":"The reader's weakest_assumption identifies exactly this junction: trusted deterministic components behave as specified while semantic fit and catalogue/curator choices sit outside enforcement, with the 21-case underspecification misses as the concrete symptom. My read converges on the same concern, sharpened to its most load-bearing form: the one place the paper empirically exercises the semantic-fit gap, the model fails 3/3 and is rescued only by frozen ordering, which the authors admit is not order-invariant. I do not recommend moving the verdict, for three reasons. First, the paper is unusually transparent: the order-dependence, the 40/40 wrong-template releases, the descriptive nature of the routing win (p=.070, template-block CI crossing zero), and the curator-owned trust perimeter are all disclosed in §4.8, Table 1, and Suppl. A/B — the claim as literally stated (\"blocked lifecycle hazards,\" \"process-compliant execution\") is supported. Second, the process-level guarantees are backed by real evidence: sealed campaigns, a 20/20 fault-injection matrix, a 188/188 verifier matrix, independent replay, and preserved No-Go predecessor records (Suppl. F), which is well above typical agent-demo practice. Third, the fix is identifiable and testable (ordering permutation plus a semantic abstention mechanism), which is precisely what CONDITIONAL is for. The one-place disagreement I might register is emphasis, not substance: the reader treats semantic abstention as one caveat among several; I read it as the single point where the central \"fail-closed\" framing most overreaches the demonstrated evidence, since every other demonstrated guarantee is about execution integrity rather than selection correctness.","tokens_in":21245,"tokens_out":2073,"duration_ms":81915,"concrete_test":"Re-run the 21-case composed campaign under permuted orderings in which each underspecified request precedes its corresponding supported request (so no contract is pre-reserved), under the same seal and scorer. Count releases for underspecified requests. If any underspecified request yields a Released trace, the fail-closed property holds only for lifecycle/process violations and the bounded-autonomy claim requires an added semantic-abstention control; if all permutations still terminate in Abstained, the order-dependence concern does not land. As a secondary check, have an external model author ~20 fresh underspecified/ambiguous phrasings from new intent cards and measure the semantic abstention rate directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim has three legs: routing superiority (39/40 vs 33/40), three hazard-blocked ERP releases, and FPR reduction via held-out confirmation. The second and third legs are well supported under stated scope. The load-bearing soft spot sits at the junction of legs one and two: the release-safety argument (Suppl. A) is explicitly conditional on trusted deterministic components and expressly excludes semantic fit (\"semantic intent and scientific optimality remain curator-governed\"). The paper's own probes quantify the resulting gap: full preflight releases all 40/40 internally consistent wrong-template substitutions (Table 9), and in the 21-case composed campaign the model mapped all three underspecified requests to already-reserved supported contracts rather than abstaining (Table 5/§B.1). Release was prevented only because the supported requests happened to run first and reserve those contracts; the authors state containment \"is not order-invariant.\" So in the demonstrated failure mode, the system would have executed and released a statistically valid, fully audited, participant-disjoint analysis of a question the user did not ask — the exact intent-alignment failure the architecture is meant to govern. Combined with the routing win being descriptive (exact paired p=.070; template-block interval −2.5 to 35.0 crosses zero) on designer-visible synthetic requests, the empirical basis for \"the model selects the requested analysis\" is thinner than the headline suggests. This qualifies rather than falsifies the claim: the paper discloses all of it, and the process-level guarantees (commitment order, single-use confirmation, evidence binding) are genuinely demonstrated by the fault matrices. But \"bounded autonomy\" currently bounds execution, not selection.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript presents CogEEGAgent, an LLM agent for cognitive-EEG analysis built on MNE-Python, in which the LLM performs bounded semantic routing (mapping natural language to a registered analysis template) while deterministic components own contract materialization, participant-disjoint discovery/confirmation, commitment, falsification, and evidence-bound release. Three prospective studies support the claims: (i) a sealed 60-request routing benchmark in which Qwen2.5-14B achieves 39/40 exact routes versus 33/40 for a matched field-weighted BM25 router, with shared preflight forcing abstention on all 20 required-abstention requests; (ii) a 21-case externally model-authored, outcome-blind campaign in which three supported ERP analyses (ERN, N400, MMN) are released with participant-disjoint confirmation and all hazard/reuse requests are blocked, though the three underspecified requests are misrouted to already-reserved contracts and contained only by lifecycle ordering; (iii) a 100,000-repetition policy stress test showing held-out confirmation reduces adaptive-search FPR from 16.5% to 4.9%. Deterministic enforcement is exercised by 20/20 and 188/188 fault-injection fixtures.","tokens_in":21653,"tokens_out":5191,"duration_ms":193077,"significance":"If the results hold, this is a useful and unusually disciplined contribution to the scientific-agents literature: the separation of semantic from scientific authority, the executable commit-before-confirmation state machine, and the participant-disjoint release path operationalize established sample-splitting policy in an agent setting. Strengths that materially raise confidence: sealed protocols with hash-bound ledgers and scorer-only gold opened after model calls stop; deterministic fault-injection matrices with exact match counts; externally authored outcome-blind requests; explicit preservation of failed predecessors (the No-Go P3/N170 pilot, §F.2); participant-disjoint confirmation with sign-flip and leave-one-out falsification; and public code. The paper is also commendably transparent about its own boundary (Table 1, §4.8, Suppl. A). The main residual risk is that the headline \"fail-closed\" framing exceeds the demonstrated enforcement scope, which is process-level rather than intent-level.","major_comments":[{"comment":"The abstract's closing claim of \"fail-closed control over inference and release\" is not matched by the paper's own semantic-misfit evidence. In the composed campaign, all three policy-designated underspecified requests were routed to already-reserved supported contracts (0/3 correct abstentions, Table 5/§B.1) and release was prevented only because the supported requests happened to run first; the authors state containment \"is not order-invariant.\" Table 9 further shows full preflight releases 40/40 internally consistent wrong-template substitutions. In the counterfactual ordering, the system would release a statistically valid, fully audited, participant-disjoint analysis of a question the user did not ask. Two requests: (a) report the counterfactual explicitly — the runtime is deterministic, so either re-run the campaign with underspecified requests ordered first or state precisely what","section":"Abstract/§5 vs §B.1 Table 5 and §4.1 Table 9"},{"comment":"The abstract states the agent \"maps language to registered analyses more accurately than a matched deterministic router\" without the qualification the body itself supplies: the exact paired test gives p=.070, the template-block sensitivity interval (−2.5 to 35.0 points) crosses zero, the corpus is designer-visible and synthetic, and implementation was finalized after the wording was visible (§4.8). The body correctly calls the advantage \"descriptive and corpus-specific\"; the abstract and contribution 1 should carry the same hedge, since routing superiority is one of the three legs of the evidence chain and will be the most-cited number (39/40 vs 33/40).","section":"§4.1 and Abstract, routing claim"},{"comment":"The positive end-to-end evidence is three supported releases (8/8, 10/10, 19/19 participants), one LLM, one frozen request ordering, and familiar pre-loaded data. This is appropriate as a demonstration, but the abstract's \"these studies establish bounded autonomy\" is stronger than n=3 semantic successes plus a 3/3 semantic-abstention failure supports. Please temper \"establish\" (e.g., \"demonstrate ... under the stated trusted-boundary and catalogue assumptions\") or add evidence (additional phrasings, a second model on the composed path, or the order-permutation sensitivity requested above).","section":"§4.2/§4.6, breadth of the positive release evidence"}],"minor_comments":[{"comment":"Row label renders as \"Best-p16.5\" — the policy name and the FPR entry appear to have lost a column separator.","section":"Table 4"},{"comment":"PCVR's injected-error detection F1 is modest overall (.442; .400–.471 per paradigm). The main text notes PCVR is \"nearly inert in this conservative-model regime\" but does not report the F1; adding it would prevent over-reading of the post-hoc audit path's detection power.","section":"§4.4 / Table 11"},{"comment":"The external authoring session showed a fallback from Claude Fable 5 to Opus 4.8 and account-level memory status was unverifiable. A sentence on why this poses negligible contamination risk (e.g., no registry identifiers or outcomes were provided) would strengthen the provenance account, which is otherwise admirably detailed.","section":"Suppl. §B.1"},{"comment":"The effect-evidence score f and its cluster-count fallback min(1, 0.5·n_clusters) are ad hoc; since τ=0.3 halts on cumulative effect change, a brief sensitivity note on these choices would help.","section":"§3.3 / Suppl. §H"},{"comment":"The Native arm's \"Rel.\" column (e.g., 12/12 released for unsupported-template requests) counts pre-guard model routes; a one-line caption clarification would avoid confusion with the post-preflight release counts elsewhere.","section":"Table 8"},{"comment":"ds003061 (Delorme 2021) appears only in the auxiliary scope section; a pointer in the main text or a datasets summary would help readers track which data enter which evaluation.","section":"§N"},{"comment":"Please state the Qwen2.5-14B serving/quantization details (or a precise pointer to the runtime manifest) in the main text, since the single-GPU feasibility claim in §4.5 depends on it.","section":"§4.1 / Suppl. §C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more self-critical than most agent papers: failed pilots are preserved, the descriptive status of the routing win is stated, and the order-dependence of the underspecified-request containment is conceded in §4.8. My major comments therefore ask for claim calibration and one cheap deterministic counterfactual rather than new science. Citation pattern includes several very recent self-citations (Hou et al. 2026; Zeng et al. 2026), all topically relevant; no action needed. Fit for the venue is good. I would accept after the abstract/conclusion scoping and the counterfactual-ordering disclosure."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is not another EEG-agent demo that confuses fluent reports with valid inference. They built a real control pattern: LLM only picks a registered route; deterministic code owns contracts, commit-before-confirm, one-shot hidden confirmation, falsification, and evidence-bound release. That separation is the contribution, and they stress-tested it harder than most agent papers do.\n\nWhat works: the sealed composed campaign actually releases three participant-disjoint ERP analyses (ERN/N400/MMN) with independent replay and blocks the listed lifecycle hazards. The fault matrices (20/20 runtime, 188/188 verifier) and the policy study (best-p FPR 16.5% → held-out 4.9%) are the load-bearing evidence. Citations are in the right neighborhood—scientific agents, holdout validity, ERP CORE, MNE—and self-cites look like prior domain work, not circular scaffolding. Code is public; seals and scorers are described in enough detail to take seriously.\n\nSoft spots, in proportion: the stress-test note is basically right. Release safety is process-closed, not intent-closed. Preflight happily releases internally consistent wrong templates (40/40), and the three underspecified campaign requests all mapped to already-reserved supported contracts; non-release depended on ordering, which they concede is not order-invariant. So a statistically clean, fully audited answer to the wrong question is still possible. The routing win over FW-BM25 (39/40 vs 33/40) is descriptive (paired p=.070; template-block CI crosses zero) on designer-visible synthetic requests. None of that kills the paper—they disclose it—but “bounded autonomy” currently bounds execution more than selection.\n\nWho it’s for: people building or reviewing scientific LLM agents, and EEG labs that want auditable automation over a fixed catalogue. Not a methods breakthrough in ERP statistics.\n\nI’d send it to referees. Engage if you care about agent trust boundaries; skim the campaign and Suppl. A if you only need the pattern.","headline":"Solid architecture paper: process-level fail-closed control is real and carefully evaluated; intent alignment is still mostly curator- and catalogue-bound, which the authors mostly admit.","tokens_in":22733,"tokens_out":530,"would_cite":true,"duration_ms":14272,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A cognitive-EEG agent can answer natural-language analysis requests while a fail-closed harness, not the model, decides what evidence may be released.","keywords":["cognitive EEG","LLM agents","scientific automation","selection-aware verification","event-related potentials","fail-closed release","MNE-Python","adaptive analysis"],"falsifier":"Run the sealed composed campaign with underspecified or wrong-template requests and check whether any report still releases without a matching committed, participant-disjoint confirmation return, or whether held-out confirmation fails to keep adaptive-search false positives near the nominal rate on the same candidate families.","tokens_in":22461,"feed_emoji":"🧠","tokens_out":908,"duration_ms":19903,"temperature":0.7,"pith_summary":"Cognitive EEG analysis is full of defensible choices—contrasts, channels, windows, tests—and fluent AI reports do not prove that the right analysis was run or that a confirmatory claim was tested without fishing. This paper presents CogEEGAgent, which lets a language model map a request onto a registered analysis while deterministic code owns contracts, data access, commitment, confirmation, and release. On a sealed routing benchmark the model routes more accurately than a matched lexical router, and shared preflight forces abstention when required. In an outcome-blind end-to-end campaign the full system releases three supported ERP analyses with participant-disjoint confirmation and blocks lifecycle hazards. A separate policy stress test shows held-out confirmation cuts false positives from uncorrected adaptive search. The point is bounded autonomy: flexible language understanding without handing scientific authority to the model.","feed_headline":"EEG agent routes language, but code alone can release claims","feed_subtitle":"Commit-then-confirm harness blocks fishing and lifecycle reuse while still answering registered ERP questions","key_machinery":"The EEG-specific scientific harness—especially the Scientific Control Plane with trace–partition–evidence binding—commits one discovery candidate before single-use participant-disjoint confirmation and authorizes release only from typed, replayed evidence. Paradigm-Conditioned Verification audits completed workflows post hoc but does not replace that prospective boundary.","core_discovery":"CogEEGAgent establishes that cognitive-EEG workflows can be automated under bounded autonomy by separating semantic authority from scientific authority. The language model only chooses a registered route; a Scientific Control Plane materializes typed contracts, commits one discovery selection before opening hidden confirmation once, and releases only evidence-bound reports. Sealed studies show higher exact routing than a matched deterministic router, release of supported analyses with blocked capability and reuse hazards, and held-out confirmation restoring false-positive control under adaptive search.","pith_inferences":["The remaining failure mode is semantic abstention under underspecification: process safety can block release while still selecting the wrong reserved contract, so intent validation is the next bottleneck.","Domains without a separable confirmation split can reuse the post-hoc audit path but cannot claim the same selection-aware guarantee.","Catalogue quality becomes a first-class scientific object: incomplete or biased registered families would make process-compliant releases still scientifically narrow.","Completion metrics that ignore successful statistical returns will systematically overrate models that narrate effects without calling tools."],"forward_implications":["Natural-language EEG questions can be answered with receipts that link contract, commitment, confirmation, and report rather than model prose alone.","Scientific agents can be scored on semantic routing, commitment order, evidence lineage, and release—not only task completion.","When full adaptive-search history cannot be trusted, held-out confirmation is a practical fail-closed substitute for uncorrected best-p selection.","The same commit-then-confirm interfaces can wrap other registered scientific executors that expose a finite catalogue and a reserved confirmation resource.","Model revisions can be swapped against fixed contracts without changing release semantics."],"fun_headline_variants":["EEG agent routes language; code alone releases claims","CogEEGAgent separates intent from evidence-bound release","LLM proposes EEG routes; control plane gates confirmation","Bounded EEG autonomy: register, commit, then confirm once","Held-out confirmation curbs adaptive search false positives"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The guarantee holds only if the frozen registry, deterministic executor, commitment ledger, confirmation opener, and scorer behave as specified, while the curator’s catalogue and the model’s semantic fit remain trusted outside automatic enforcement.","fun_headline_variants_meta":{"raw":{"variants":["EEG agent routes language; code alone releases claims","CogEEGAgent separates intent from evidence-bound release","LLM proposes EEG routes; control plane gates confirmation","Bounded EEG autonomy: register, commit, then confirm once","Held-out confirmation curbs adaptive search false positives"]},"model":"grok-4.5","effort":"low","cost_usd":0.004483,"raw_usage":{"total_tokens":1361,"prompt_tokens":812,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":44828000,"prompt_tokens_details":{"text_tokens":812,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":491,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":812,"tokens_out":58,"duration_ms":7869,"temperature":1.0,"reasoning_tokens":491,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T02:36:55.554745+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the sealed composed campaign with underspecified or wrong-template requests and check whether any report still releases without a matching committed, participant-disjoint confirmation return, or whether held-out confirmation fails to keep adaptive-search false positives near the nominal rate on the same candidate families.","supporting_citations":[],"review_version":1}