{"id":"f7c0b398-f2c2-40fe-8614-b323f887d95e","arxiv_id":"2608.02646","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Cross-anesthetic ECoG state decoding fails at the decision threshold, not the representation, at least within-subject for five mouse anesthetics.","lead":"A brain-signal decoder that separates awake from anesthetized mice keeps the correct ordering for every drug it has never seen, including ketamine, but its fixed cutoff calls ketamine anesthesia awake. The paper shows the failure is a misplaced decision threshold, not a missing neural representation, and a pre-drug baseline recording fixes it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Session-level AUROC inflates the 'representation transfers' claim: ketamine epoch-level AUROC is only 0.68, so at the decision timescale the failure is not purely threshold.","rationale":"The paper makes a clearly scoped, internally consistent empirical case at the session level: band-power ranks awake versus anesthetized on every held-out drug, and the fixed-threshold balanced accuracy collapses for ketamine while the ranking does not. The authors disclose the epoch-level AUROC (0.68) and state that all inferential statistics are session-level. However, the headline claim—'fails at the decision threshold, not the representation'—is materially weaker at the epoch level, which is the timescale at which a state decoder actually operates. The session-level AUROC is obtained by averaging up to 186 overlapping, autocorrelated epochs; such averaging can dramatically inflate AUROC when each session is internally homogeneous and the number of sessions is tiny (5 effect, 15 total from 3 mice). The paper's per-session z-scoring removes between-session mean differences in band power, so the high session-level AUROC must arise from higher-order statistics or session-level confounds, and its relevance to epoch-level decoding is unclear. If the epoch-level AUROC were truly 0.68, then even an optimally placed threshold could not yield balanced accuracy much above 0.65–0.70; the reported 0.85 improvement is a session-level artifact. This does not invalidate the session-level dissociation between ranking and threshold, but it means the 'representation transfers' claim is not established at the decision-relevant timescale. The reader's weakest_assumption identified exactly this epoch-versus-session gap, and the paper honestly reports the limitation. The appropriate disposition remains conditional: the central claim should be accepted only if the authors either demonstrate epoch-level calibration performance or explicitly scope the claim to session-averaged decoding. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":11906,"tokens_out":7397,"duration_ms":66000,"concrete_test":"Recompute ketamine balanced accuracy with the baseline-anchored threshold applied to each 2-s epoch (not session-averaged scores), and report the epoch-level AUROC with a mouse-cluster bootstrap CI (resample mice, then sessions, then epochs). If the epoch-level calibrated balanced accuracy is ≤0.70 or the epoch-level AUROC CI includes 0.5, the failure is not purely at the threshold at the decision timescale, and the central claim needs to be scoped to session-averaged decoding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ketamine failure is at the threshold, not the representation, rests on session-level AUROC ≥0.96, but the paper's own descriptive epoch-level AUROC for ketamine is 0.68. Decisions in the decoder are made on 2-s epochs; the session-level figure averages up to 186 overlapping, autocorrelated epochs, which can inflate AUROC whenever sessions are internally homogeneous and the session count is small (5 effect, 15 total sessions from 3 mice). With per-session z-scoring, between-session mean differences in band power are removed; remaining session separation may reflect higher-order statistics or non-neural session-level confounds. At the epoch level, an AUROC of 0.68 means a large fraction of anesthetized epochs are not separable from awake epochs, so a threshold shift cannot recover strong epoch-level discrimination. The paper reports fixed-threshold balanced accuracy only at session level (0.500→0.850 after calibration); if the baseline-anchored threshold were applied to individual epochs, the achievable balanced accuracy is bounded by the epoch-level separation (AUROC 0.68), likely ≈0.65–0.70, not 0.85. Thus the 'representation transfers' claim is unverified at the timescale where the decoder actually operates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper dissociates representation from decision-threshold failures in leave-one-anesthetic-out (LOAO) decoding of awake versus anesthetized states from mouse ECoG. A spatially blind band-power decoder ranks each held-out drug at session-level AUROC ≥0.96 (ketamine 0.980), yet at the inherited 0.5 threshold it labels nearly all ketamine effect sessions awake (balanced accuracy 0.500). A causal, label-free threshold set to the 95th percentile of each drug's pre-induction awake scores raises ketamine session-level balanced accuracy to 0.850 and, on average, outperforms Riemannian domain adaptation, which is net-negative. All inferential statistics are session-level with mouse-cluster bootstrapping for the ketamine fold; the epoch-level AUROC for ketamine is reported descriptively as 0.680. The authors explicitly scope the result to within-subject cross-drug transfer, to the ≤250 Hz LFP band, and to the small ketamine sample (5 effect sessions from 3 mice).","tokens_in":12127,"tokens_out":8304,"duration_ms":73869,"significance":"If the dissociation holds, the paper makes a valuable methodological point: in transfer settings, a ranking metric (AUROC) and a thresholded metric (balanced accuracy) can diverge, and calibration should be reported and corrected separately from representation quality. The baseline-anchored threshold is simple, causal, and label-free, and it is contrasted honestly with a non-causal Riemannian domain-adaptation upper bound. The authors are commendably transparent: q=0.95 was fixed a priori, RA is labeled a non-causal upper bound throughout, the ketamine power ceiling is stated repeatedly, and the epoch-level AUROC is disclosed. However, the headline conclusion—that the failure is 'not the representation'—is stronger than what the session-level evidence supports, because the decoder's own decision timescale is the epoch, and the paper's reported epoch-level separation is only 0.680. This mismatch must be addressed before the central claim is publishable at face value.","major_comments":[{"comment":"The central claim that the ketamine representation transfers is supported only by session-level AUROC (0.980; 95% CI [0.885, 1.000] session-level, [0.821, 1.000] mouse-cluster), while the paper's own epoch-level AUROC for ketamine is 0.680, described as descriptive. The BP decoder emits per-epoch scores and the decision threshold is applied to those scores, so the ranking at the actual decision timescale is 0.68, not 0.98. With an ROC AUC of 0.68, no threshold shift can yield epoch-level balanced accuracy much above roughly 0.65–0.70; therefore the reported 0.850 calibrated balanced accuracy in Table II must reflect session-level aggregation of up to 186 overlapping, autocorrelated epochs. The dissociation 'ranking preserved, threshold mis-scaled' is thus demonstrated only for session-averaged scores, not for the per-epoch decisions on which the decoder operates. Please report the baseline-anchored threshold's balanced accuracy at the epoch level for ketamine (and ideally for all drugs), or explicitly redefine the decision unit as the session average and adjust the title and abstract accordingly. Without this, the conclusion that the failure is 'not the representation' is not established at the timescale of actual decisions. The same caveat applies to the permutation test: the significant ranking result (p=0.0025) is computed on session-level AUROC and cannot establish that the per-epoch representation is intact.","section":"III.A and Methods F"},{"comment":"The abstract states that baseline calibration 'fixes ketamine (balanced accuracy 0.50 to 0.85)', but the mouse-cluster 95% CI for this estimate is [0.510, 1.000], which includes chance. Although the text reports this honestly in Section III.D, the abstract and the phrase 'fixes ketamine' overstate the strength of the correction. Please either add the CI to the abstract/conclusion or soften the wording to indicate that the point improvement is not statistically tight given the three-mouse, five-effect-session ceiling.","section":"Results D and Abstract"}],"minor_comments":[{"comment":"The second author name appears as 'Qianwei zhou'; the surname should be capitalized as 'Qianwei Zhou'.","section":"Title page"},{"comment":"Several references contain placeholder text 'pMID: ...' and 'verify volume/pages at typesetting'; these must be resolved before publication.","section":"References [6], [9], [24]"},{"comment":"The abbreviation 'OAS' (Oracle Approximating Shrinkage) is used without definition or citation; please define it on first use.","section":"Methods D"},{"comment":"The metric 'effect-called-awake' is introduced with values '1.00→0.20' without a definition; please state explicitly that it is the fraction of anesthetized sessions whose mean epoch score falls below the threshold.","section":"Results D"},{"comment":"The sentence 'the lower bound falls ~0.06 but remains far from chance' is imprecise; it would be clearer to state that the lower bound drops from 0.885 to 0.821 under mouse-cluster resampling.","section":"Results A"},{"comment":"There is a typo in 'for a statedecoder it should not be corrected away'; it should be 'for a state decoder'.","section":"Section IV.C"}],"recommendation":"major_revision","confidential_remarks":"The paper has a strong, well-scoped core result, but the title and abstract currently overclaim relative to the session-level evidence. The authors are transparent about the epoch-level AUROC of 0.68, which actually undercuts the central dissociation at the decision timescale; I recommend asking for an epoch-level calibration analysis or an explicit redefinition of the decision unit. The data availability statement ('can be made available to qualified researchers') is vague; the journal may want a concrete repository for code and derived tables. Overall, the manuscript is worth a major-revision opportunity rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely useful: it stops treating cross-drug decoder failure as a single accuracy number and separates the ranking question (does the anesthetic state still order correctly?) from the threshold question (does the inherited boundary sit in the right place?). In a controlled mouse ECoG preparation with five mechanistically distinct anesthetics and strict leave-one-drug-out, the band-power representation ranks every held-out drug at session AUROC ≥0.96, including ketamine, while fixed-threshold balanced accuracy collapses to chance. That dissociation, backed by a permutation test that is significant for ranking but not for thresholded accuracy, is the core contribution. The baseline-anchored threshold—using only pre-induction awake scores, no anesthetized labels—is a simple, deployable correction that dominates Riemannian domain adaptation in their comparison. The authors are also admirably scoped: they call the ketamine result within-subject, label RA as a non-causal upper bound, and report the wide mouse-cluster confidence intervals. Now the soft spots, in proportion. The biggest is exactly what the stress-test note flags: the representation claim rests on session-level AUROC, but the decoder operates on 2-s epochs, and the ketamine epoch-level AUROC is only 0.68. That means the strong 'representation transfers' statement is at a coarser timescale than the actual decision. The paper is honest about this—epoch-level is reported only descriptively—but the abstract and framing let the session-level number carry the headline. If a reader thinks the baseline calibration gives 0.85 balanced accuracy per epoch, that is not supported; the epoch-level bound is closer to 0.65–0.70. So the deployable-fix claim is overstated unless clarified to session-level decision-making, which is not how such a decoder would run. Second, the ketamine evidence is thin: five effect sessions, three mice, and no novel-subject ketamine testing. The cluster bootstrap is appropriate and the CI touches 0.51 for the calibration gain, so the magnitude of the fix is uncertain. Third, no public code or data, which matters for a result this counter to the field's default reflex. None of this sinks the paper. The core dissociation is a real empirical observation, the statistics are handled carefully (session-level inference, cluster bootstrap, permutation tests), and the limitations are disclosed rather than buried. I would send this to peer review without hesitation, but I would ask the authors to either provide epoch-level balanced accuracy for the baseline-calibrated threshold or explicitly state that the reported 0.85 is a session-level aggregate. The paper is for anyone working on transfer learning in brain-state decoding, anesthesia monitoring, or calibration under distribution shift. It deserves a serious referee, and the right revision will make the timescale of the claim unambiguous.","headline":"Clean separation of ranking from threshold in cross-drug decoding, with an honest but important caveat that the 'representation transfers' claim is session-level while real decisions happen per epoch.","tokens_in":844,"tokens_out":769,"would_cite":false,"duration_ms":27186,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cross-drug anesthesia decoding fails at the threshold, not at the representation.","keywords":["anesthesia depth monitoring","electrocorticography","calibration","domain adaptation","transfer learning","ketamine","AUROC","decision threshold"],"falsifier":"Record many more ketamine anesthesia sessions (at least twenty effect sessions across at least five mice), score them with the same band-power decoder trained on the other four drugs, and compute AUROC at the epoch level and at the session level with confidence intervals that account for within-mouse and within-session correlation; if the session-level AUROC confidence interval includes 0.5, or the epoch-level AUROC stays near 0.5, the claim that the representation transfers is refuted.","tokens_in":11669,"feed_emoji":"🧠","tokens_out":7794,"duration_ms":67204,"temperature":0.7,"pith_summary":"This paper separates two failure modes that a single accuracy number fuses: a decoder can lose the neural distinction under a new drug, or keep the distinction intact and simply place its decision boundary wrongly. The paper argues that, in mouse ECoG with leave-one-anesthetic-out evaluation, the distinction transfers across five mechanistically distinct anesthetics while the inherited boundary does not. A spatially blind band-power model ranks awake versus anesthetized on every held-out drug, including ketamine, yet at the threshold trained on the other drugs it calls all ketamine anesthetized sessions awake. The practical stake is that the right fix is recalibrating the decision point with a pre-induction baseline, not enriching features or applying domain adaptation. That matters for depth-of-anesthesia monitoring, where ketamine is the notorious transfer failure and the standard remedy points at the wrong quantity.","feed_headline":"Cross-drug anesthesia decoding fails at threshold, not representation","feed_subtitle":"Held-out ketamine is still ranked correctly; only the decision point is misplaced, and a baseline fix repairs it.","key_machinery":"The central decomposition is the pair of metrics applied to the same decoder output: session-level AUROC, which is invariant to where the decision boundary sits, and balanced accuracy at a fixed threshold, which is not. Around that pair the paper builds leave-one-anesthetic-out evaluation across five drugs, a baseline-anchored threshold set at the 95th percentile of the held-out drug's own pre-induction awake scores, and mouse-level cluster bootstrapping to hedge the small ketamine sample. The band-power decoder is deliberately simple, using eight log band-power features pooled across channels, so that any transfer can be attributed to the signal rather than spatial structure. Riemannian domain adaptation is included as the standard field remedy and is evaluated as a non-causal upper bound, meaning even its best-case numbers cannot be achieved online.","core_discovery":"Under leave-one-anesthetic-out evaluation on mouse ECoG, a spatially blind band-power decoder ranks awake versus anesthetized correctly on every held-out drug, including ketamine (session AUROC 0.980; mouse-cluster CI [0.821, 1.000]), yet at the decision threshold inherited from the other four drugs it labels all ketamine effect sessions as awake, giving a balanced accuracy of 0.500. Across three representations the ketamine ranking stays nearly invariant (session AUROC 0.98, 0.94, 0.96) while thresholded accuracy swings from 0.500 to 0.850, and a permutation test is significant for ranking (p = 0.0025) but not for fixed-threshold accuracy (p = 0.3795). The paper concludes that the neural representation of anesthetic state transfers across drugs and that the failure is confined to calibration. A label-free threshold anchored to the subject's own pre-induction baseline repairs ketamine (balanced accuracy 0.850) and outperforms Riemannian domain adaptation, which is net-negative on average. The authors scope the ketamine finding to within-subject cross-drug transfer, based on five effect sessions from three mice, and explicitly decline to claim population-level transfer across subjects.","pith_inferences":["If the same decomposition holds in human scalp EEG, the fastest route to a ketamine-tolerant depth monitor may be recalibrating the decision boundary from a short pre-induction baseline rather than redesigning the classifier; this is testable now on existing human ketamine datasets.","The gap between epoch-level AUROC (0.68) and session-level AUROC (0.98) implies that averaging many overlapping epochs inflates the apparent transfer, so a fair deployment test should use non-overlapping or properly de-correlated windows.","The fact that the same baseline information succeeds as a threshold shift but fails as feature normalization suggests that for covariate shifts with a known generator, the locus of the fix may matter more than the information content; analogous tests with sleep or sedation data could reveal whether threshold misplacement is the generic failure mode.","The paper's high-frequency ablation (ketamine AUROC rising from 0.767 to 0.867 when bands extend to 200 Hz) predicts that re-acquiring wideband ECoG above 250 Hz would further separate ketamine's anesthetized and awake states, directly testing the proposed high-frequency mechanism."],"forward_implications":["Cross-drug transfer of a state decoder should be reported with at least two metrics, a threshold-invariant ranking metric and a thresholded accuracy, because they can disagree completely in the same fold.","Riemannian domain adaptation should not be assumed to help cross-drug transfer: in this preparation it is net-negative on average and buys the ketamine fold only by damaging dexmedetomidine and isoflurane.","A threshold anchored to the subject's pre-induction awake baseline, using no anesthetized labels and no new model, is a deployable correction that raises mean balanced accuracy from 0.813 to 0.858 and fixes ketamine from 0.500 to 0.850.","Applying the same baseline information at the feature level and retraining does not fix ketamine (balanced accuracy stays 0.500), so the locus of the correction matters, not just the information content.","The validity of baseline anchoring is predictable: it helps when the pre-drug and post-drug awake spectra match and hurts when they differ, with base-versus-post separability rank-ordering the gains across the five drugs (Spearman rho = -1.000, n = 5, suggestive)."],"supporting_citations":[{"why":"Documents ketamine's EEG signature of high-frequency power without sustained delta, the biological generator of the score shift that leaves ranking intact but moves the boundary.","marker":"[5]"},{"why":"Shows the only prior demonstrated cross-drug transfer is GABAergic and explicitly excludes ketamine, setting the baseline the paper's leave-one-drug-out result must extend.","marker":"[13]"},{"why":"Supplies the Riemannian recentering method evaluated as the standard domain-adaptation remedy and shown to be net-negative here.","marker":"[17]"},{"why":"Defines AUROC as a threshold-independent ranking metric, the tool that separates representation from calibration in the analysis.","marker":"[19]"},{"why":"Provides independent cross-modal evidence that surgical-plane ketamine cortical dynamics can resemble quiet wakefulness, supporting the biological interpretation of ranking ketamine as awake.","marker":"[25]"}],"fun_headline_variants":["Cross-drug anesthesia decoding fails at threshold, not representation","Right ranking, wrong cutoff: cross-drug anesthesia decoding","Ketamine decoding fixed by baseline threshold, not adaptation","Representation transfers across anesthetics; threshold does not"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that ketamine's brain signal still separates from the awake state rests on only five anesthetized ketamine sessions from three mice, with the high ranking coming partly from averaging many overlapping time windows; if a larger or properly de-correlated sample does not keep the ranking above chance, the central claim fails.","fun_headline_variants_meta":{"raw":{"variants":["Cross-drug anesthesia decoding fails at threshold, not representation","Right ranking, wrong cutoff: cross-drug anesthesia decoding","Ketamine decoding fixed by baseline threshold, not adaptation","Representation transfers across anesthetics; threshold does not"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000587,"raw_usage":{"total_tokens":2857,"prompt_tokens":1148,"completion_tokens":1709,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":764,"completion_tokens_details":{"reasoning_tokens":1646}},"tokens_in":764,"tokens_out":1709,"duration_ms":12780,"temperature":1.0,"reasoning_tokens":1646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:21:11.879319+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record many more ketamine anesthesia sessions (at least twenty effect sessions across at least five mice), score them with the same band-power decoder trained on the other four drugs, and compute AUROC at the epoch level and at the session level with confidence intervals that account for within-mouse and within-session correlation; if the session-level AUROC confidence interval includes 0.5, or the epoch-level AUROC stays near 0.5, the claim that the representation transfers is refuted.","supporting_citations":[{"cited_title":"Electroencephalogram signatures of ketamine anesthesia-induced unconsciousness,","cited_arxiv_id":null,"evidence_quote":"Documents ketamine's EEG signature of high-frequency power without sustained delta, the biological generator of the score shift that leaves ranking intact but moves the boundary."},{"cited_title":"Machine learning of EEG spectra classifies unconsciousness during GABAergic anesthesia,","cited_arxiv_id":null,"evidence_quote":"Shows the only prior demonstrated cross-drug transfer is GABAergic and explicitly excludes ketamine, setting the baseline the paper's leave-one-drug-out result must extend."},{"cited_title":"Transfer learning: A Riemannian geometry framework with applications to brain- computer interfaces,","cited_arxiv_id":null,"evidence_quote":"Supplies the Riemannian recentering method evaluated as the standard domain-adaptation remedy and shown to be net-negative here."},{"cited_title":"An introduction to ROC analysis,","cited_arxiv_id":null,"evidence_quote":"Defines AUROC as a threshold-independent ranking metric, the tool that separates representation from calibration in the analysis."},{"cited_title":"Existence of multiple transitions of the critical state due to anesthetics,","cited_arxiv_id":null,"evidence_quote":"Provides independent cross-modal evidence that surgical-plane ketamine cortical dynamics can resemble quiet wakefulness, supporting the biological interpretation of ranking ketamine as awake."}],"review_version":1}