REVIEW 2 major objections 6 minor 25 references
Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation
T0 review · 2 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Cross-drug anesthesia decoding fails at the threshold, not at the representation.
desk verdict Clean separation of ranking from threshold in cross-drug decoding, with an honest but important caveat that the 'representation transfers' claim is session-level while real decisions happen per epoch. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central decomposition is the pair of metrics applied to the same decoder output: session-level AUROC, which is invariant to where the decision boundary sits, and balanced accuracy at a fixed threshold, which is not. Around that pair the paper builds leave-one-anesthetic-out evaluation across five drugs, a baseline-anchored threshold set at the 95th percentile of the held-out drug's own pre-induction awake scores, and mouse-level cluster bootstrapping to hedge the small ketamine sample. The band-power decoder is deliberately simple, using eight log band-power features pooled across channels, so that any transfer can be attributed to the signal rather than spatial structure. Riemannian domain adaptation is included as the standard field remedy and is evaluated as a non-causal upper bound, meaning even its best-case numbers cannot be achieved online.
What would settle it
Record many more ketamine anesthesia sessions (at least twenty effect sessions across at least five mice), score them with the same band-power decoder trained on the other four drugs, and compute AUROC at the epoch level and at the session level with confidence intervals that account for within-mouse and within-session correlation; if the session-level AUROC confidence interval includes 0.5, or the epoch-level AUROC stays near 0.5, the claim that the representation transfers is refuted.
Extended reading notes
Core claim
Under leave-one-anesthetic-out evaluation on mouse ECoG, a spatially blind band-power decoder ranks awake versus anesthetized correctly on every held-out drug, including ketamine (session AUROC 0.980; mouse-cluster CI [0.821, 1.000]), yet at the decision threshold inherited from the other four drugs it labels all ketamine effect sessions as awake, giving a balanced accuracy of 0.500. Across three representations the ketamine ranking stays nearly invariant (session AUROC 0.98, 0.94, 0.96) while thresholded accuracy swings from 0.500 to 0.850, and a permutation test is significant for ranking (p = 0.0025) but not for fixed-threshold accuracy (p = 0.3795). The paper concludes that the neural representation of anesthetic state transfers across drugs and that the failure is confined to calibration. A label-free threshold anchored to the subject's own pre-induction baseline repairs ketamine (balanced accuracy 0.850) and outperforms Riemannian domain adaptation, which is net-negative on average. The authors scope the ketamine finding to within-subject cross-drug transfer, based on five effect sessions from three mice, and explicitly decline to claim population-level transfer across subjects.
Load-bearing premise
The conclusion that ketamine's brain signal still separates from the awake state rests on only five anesthetized ketamine sessions from three mice, with the high ranking coming partly from averaging many overlapping time windows; if a larger or properly de-correlated sample does not keep the ranking above chance, the central claim fails.
Editorial extensions
If this is right
- Cross-drug transfer of a state decoder should be reported with at least two metrics, a threshold-invariant ranking metric and a thresholded accuracy, because they can disagree completely in the same fold.
- Riemannian domain adaptation should not be assumed to help cross-drug transfer: in this preparation it is net-negative on average and buys the ketamine fold only by damaging dexmedetomidine and isoflurane.
- A threshold anchored to the subject's pre-induction awake baseline, using no anesthetized labels and no new model, is a deployable correction that raises mean balanced accuracy from 0.813 to 0.858 and fixes ketamine from 0.500 to 0.850.
- Applying the same baseline information at the feature level and retraining does not fix ketamine (balanced accuracy stays 0.500), so the locus of the correction matters, not just the information content.
- The validity of baseline anchoring is predictable: it helps when the pre-drug and post-drug awake spectra match and hurts when they differ, with base-versus-post separability rank-ordering the gains across the five drugs (Spearman rho = -1.000, n = 5, suggestive).
Reading between the lines
- If the same decomposition holds in human scalp EEG, the fastest route to a ketamine-tolerant depth monitor may be recalibrating the decision boundary from a short pre-induction baseline rather than redesigning the classifier; this is testable now on existing human ketamine datasets.
- The gap between epoch-level AUROC (0.68) and session-level AUROC (0.98) implies that averaging many overlapping epochs inflates the apparent transfer, so a fair deployment test should use non-overlapping or properly de-correlated windows.
- The fact that the same baseline information succeeds as a threshold shift but fails as feature normalization suggests that for covariate shifts with a known generator, the locus of the fix may matter more than the information content; analogous tests with sleep or sedation data could reveal whether threshold misplacement is the generic failure mode.
- The paper's high-frequency ablation (ketamine AUROC rising from 0.767 to 0.867 when bands extend to 200 Hz) predicts that re-acquiring wideband ECoG above 250 Hz would further separate ketamine's anesthetized and awake states, directly testing the proposed high-frequency mechanism.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper dissociates representation from decision-threshold failures in leave-one-anesthetic-out (LOAO) decoding of awake versus anesthetized states from mouse ECoG. A spatially blind band-power decoder ranks each held-out drug at session-level AUROC ≥0.96 (ketamine 0.980), yet at the inherited 0.5 threshold it labels nearly all ketamine effect sessions awake (balanced accuracy 0.500). A causal, label-free threshold set to the 95th percentile of each drug's pre-induction awake scores raises ketamine session-level balanced accuracy to 0.850 and, on average, outperforms Riemannian domain adaptation, which is net-negative. All inferential statistics are session-level with mouse-cluster bootstrapping for the ketamine fold; the epoch-level AUROC for ketamine is reported descriptively as 0.680. The authors explicitly scope the result to within-subject cross-drug transfer, to the ≤250 Hz LFP band, and to the small ketamine sample (5 effect sessions from 3 mice).
Significance. If the dissociation holds, the paper makes a valuable methodological point: in transfer settings, a ranking metric (AUROC) and a thresholded metric (balanced accuracy) can diverge, and calibration should be reported and corrected separately from representation quality. The baseline-anchored threshold is simple, causal, and label-free, and it is contrasted honestly with a non-causal Riemannian domain-adaptation upper bound. The authors are commendably transparent: q=0.95 was fixed a priori, RA is labeled a non-causal upper bound throughout, the ketamine power ceiling is stated repeatedly, and the epoch-level AUROC is disclosed. However, the headline conclusion—that the failure is 'not the representation'—is stronger than what the session-level evidence supports, because the decoder's own decision timescale is the epoch, and the paper's reported epoch-level separation is only 0.680. This mismatch must be addressed before the central claim is publishable at face value.
major comments (2)
- [III.A and Methods F] The central claim that the ketamine representation transfers is supported only by session-level AUROC (0.980; 95% CI [0.885, 1.000] session-level, [0.821, 1.000] mouse-cluster), while the paper's own epoch-level AUROC for ketamine is 0.680, described as descriptive. The BP decoder emits per-epoch scores and the decision threshold is applied to those scores, so the ranking at the actual decision timescale is 0.68, not 0.98. With an ROC AUC of 0.68, no threshold shift can yield epoch-level balanced accuracy much above roughly 0.65–0.70; therefore the reported 0.850 calibrated balanced accuracy in Table II must reflect session-level aggregation of up to 186 overlapping, autocorrelated epochs. The dissociation 'ranking preserved, threshold mis-scaled' is thus demonstrated only for session-averaged scores, not for the per-epoch decisions on which the decoder operates. Please report the baseline-anchored threshold's balanced accuracy at the epoch level for ketamine (and ideally for all drugs), or explicitly redefine the decision unit as the session average and adjust the title and abstract accordingly. Without this, the conclusion that the failure is 'not the representation' is not established at the timescale of actual decisions. The same caveat applies to the permutation test: the significant ranking result (p=0.0025) is computed on session-level AUROC and cannot establish that the per-epoch representation is intact.
- [Results D and Abstract] The abstract states that baseline calibration 'fixes ketamine (balanced accuracy 0.50 to 0.85)', but the mouse-cluster 95% CI for this estimate is [0.510, 1.000], which includes chance. Although the text reports this honestly in Section III.D, the abstract and the phrase 'fixes ketamine' overstate the strength of the correction. Please either add the CI to the abstract/conclusion or soften the wording to indicate that the point improvement is not statistically tight given the three-mouse, five-effect-session ceiling.
minor comments (6)
- [Title page] The second author name appears as 'Qianwei zhou'; the surname should be capitalized as 'Qianwei Zhou'.
- [References [6], [9], [24]] Several references contain placeholder text 'pMID: ...' and 'verify volume/pages at typesetting'; these must be resolved before publication.
- [Methods D] The abbreviation 'OAS' (Oracle Approximating Shrinkage) is used without definition or citation; please define it on first use.
- [Results D] The metric 'effect-called-awake' is introduced with values '1.00→0.20' without a definition; please state explicitly that it is the fraction of anesthetized sessions whose mean epoch score falls below the threshold.
- [Results A] The sentence 'the lower bound falls ~0.06 but remains far from chance' is imprecise; it would be clearer to state that the lower bound drops from 0.885 to 0.821 under mouse-cluster resampling.
- [Section IV.C] There is a typo in 'for a statedecoder it should not be corrected away'; it should be 'for a state decoder'.
Circularity Check
No circularity: held-out drug evaluation, a label-free baseline threshold, and explicit non-causal bounds keep the derivation self-contained.
full rationale
The paper's central claims are evaluated on held-out anesthetics: the band-power decoder is trained on four drugs and tested on the fifth, so the reported session AUROC values (ketamine 0.980, all drugs >=0.96) are genuine out-of-drug rankings, not refits. The baseline-anchored threshold is set exclusively from the held-out drug's own pre-induction awake (base) sessions, so it uses no anesthetized labels and is causal; the q=0.95 value is declared a priori with a full q-sweep reported. The permutation tests shuffle session labels, and the representation/alignment contrast (BP, RN, RA) is an evaluation of three fixed decoders rather than a fitted parameter renamed as a prediction. RA is explicitly labeled a non-causal upper bound and is used as a comparison ceiling, not as evidence for the main claim. The epoch-level AUROC of 0.68 versus the session-level 0.98 is a statistical validity concern about temporal averaging and small session count, not a circularity: the session-level metric is not defined in terms of the threshold claim, and the authors openly state the epoch-level result is descriptive. No load-bearing step reduces to its own inputs, and there is no reliance on author self-citation or an imported uniqueness theorem. The within-subject scope is disclosed, and the wide mouse-cluster confidence intervals are honestly reported, so the derivation chain remains externally anchored to the held-out data.
Assumptions & free parameters
free parameters (2)
- baseline threshold percentile q =
0.95
- per-session epoch cap =
186 epochs
assumptions (5)
- domain assumption Protocol phase labels (base, effect, post) are ground truth for awake versus anesthetized
- domain assumption The Intan LFP band (<=250 Hz) is sufficient to rank awake versus anesthetized for all five anesthetics
- domain assumption Session-level aggregation over up to 186 autocorrelated epochs yields valid AUROC estimates
- domain assumption Mouse-level cluster bootstrap is an appropriate correction for within-mouse dependence in the ketamine fold
- standard math Riemannian geometry and tangent-space projection are valid for covariance matrices of ECoG epochs
Cite this review
Pith. "Pith review of Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation." pith.science (2026). https://pith.science/paper/AXZNTWEL
@misc{pith2026260802646,
author = {Pith},
title = {Pith review of: Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation},
year = {2026},
howpublished = {\url{https://pith.science/paper/AXZNTWEL}},
note = {Machine review of arXiv:2608.02646}
}
read the original abstract
Decoders of anesthetic state from cortical activity fail across drug classes, most notoriously ketamine, but reported accuracy cannot say whether the neural representation or only the decision threshold has failed; we separate the two in a controlled preparation with ground-truth labels. We decoded awake versus anesthetized from mouse electrocorticography (the 250-Hz-bandlimited local field potential, sampled at 1875 Hz) under leave-one-anesthetic-out evaluation across five mechanistically distinct anesthetics (isoflurane, dexmedetomidine, ketamine, propofol, and midazolam), comparing a spatially blind band-power decoder, covariance/Riemannian representations, and Riemannian domain adaptation, with all statistics at the session level and mouse-level cluster bootstrapping for the ketamine fold. The representation transfers: band-power ranks awake versus anesthetized at a session AUROC of at least 0.96 on every held-out drug, ketamine included (0.980, cluster confidence interval 0.821 to 1.000). The failure is confined to the threshold: across three representations the ketamine ranking is near-invariant while its balanced accuracy swings from chance to high, and a permutation test is significant for ranking (p = 0.0025) but not for fixed-threshold accuracy (p = 0.3795). Riemannian domain adaptation is net-negative. A causal, label-free threshold anchored to the subject's own pre-induction baseline fixes ketamine (balanced accuracy 0.50 to 0.85) and dominates domain adaptation. Because the ketamine test sessions come from three mice that also contribute training drugs, this is within-subject cross-drug transfer; we do not claim population-level transfer across subjects. In cross-drug state decoding the actionable failure is calibration, not representation.
Figures
Reference graph
Works this paper leans on
-
[1]
A. A. Dahaba, “Different conditions that could result in the bispectral index indicating an incorrect hypnotic state,”Anesth. Analg., vol. 101, no. 3, pp. 765–773, 2005
work page 2005
-
[2]
Clinical electroencephalography for anesthesiologists: Part I: Background and basic signatures,
P. L. Purdon, A. Sampson, K. J. Pavone, and E. N. Brown, “Clinical electroencephalography for anesthesiologists: Part I: Background and basic signatures,”Anesthesiology, vol. 123, no. 4, pp. 937–960, 2015
work page 2015
-
[3]
Electroencephalogram signatures of loss and recovery of consciousness from propofol,
P. L. Purdon, E. T. Pierce, E. A. Mukamelet al., “Electroencephalogram signatures of loss and recovery of consciousness from propofol,”Proc. Natl. Acad. Sci. USA, vol. 110, no. 12, pp. E1142–E1151, 2013
work page 2013
-
[4]
O. Akeju, K. J. Pavone, M. B. Westoveret al., “A comparison of propofol- and dexmedetomidine-induced electroencephalogram dynam- ics using spectral and coherence analysis,”Anesthesiology, vol. 121, no. 5, pp. 978–989, 2014
work page 2014
-
[5]
Electroencephalogram signatures of ketamine anesthesia-induced unconsciousness,
O. Akeju, A. H. Song, A. E. Hamiloset al., “Electroencephalogram signatures of ketamine anesthesia-induced unconsciousness,”Clin. Neu- rophysiol., vol. 127, no. 6, pp. 2414–2422, 2016
work page 2016
-
[6]
Monitoring the depth of anesthesia using entropy features and an artificial neural network,
R. Shalbaf, H. Behnam, and H. Jelveh Moghadam, “Monitoring the depth of anesthesia using entropy features and an artificial neural network,”J. Neurosci. Methods, 2013, pMID: 23567809; verify vol- ume/pages at typesetting
work page 2013
-
[7]
Use of multiple EEG features and artificial neural network to monitor the depth of anesthesia,
Y . Gu, Z. Liang, and S. Hagihira, “Use of multiple EEG features and artificial neural network to monitor the depth of anesthesia,”Sensors, vol. 19, no. 11, p. 2499, 2019
work page 2019
-
[8]
W. Saadeh, F. H. Khan, and M. A. B. Altaf, “Design and implementation of a machine learning based EEG processor for accurate estimation of depth of anesthesia,”IEEE Trans. Biomed. Circuits Syst., vol. 13, no. 4, pp. 658–669, 2019
work page 2019
Show all 25 references
-
[9]
Electroencephalogram based detection of deep sedation in ICU patients using atomic decom- position,
S. B. Nagaraj, L. M. McClain, D. W. Zhouet al., “Electroencephalogram based detection of deep sedation in ICU patients using atomic decom- position,”IEEE Trans. Biomed. Eng., 2018, pMID: 29993386; verify volume/pages/DOI at typesetting
2018
-
[10]
Novel drug- independent sedation level estimation based on machine learning of quantitative frontal electroencephalogram features in healthy volun- teers,
S. M. Ramaswamy, M. H. Kuizenga, M. A. S. Weerink, H. E. M. Vereecke, M. M. R. F. Struys, and S. B. Nagaraj, “Novel drug- independent sedation level estimation based on machine learning of quantitative frontal electroencephalogram features in healthy volun- teers,”Br . J. Anae...
2019
-
[11]
Frontal elec- troencephalogram based drug, sex, and age independent sedation level prediction using non-linear machine learning algorithms,
S. M. Ramaswamy, M. H. Kuizenga, M. A. S. Weerink, H. E. M. Vereecke, M. M. R. F. Struys, and S. Belur Nagaraj, “Frontal elec- troencephalogram based drug, sex, and age independent sedation level prediction using non-linear machine learning algorithms,”J. Clin. Monit. Comput.,...
2022
-
[12]
Inference of brain states under anesthesia with meta learning based deep learning models,
Q. Wang, F. Liu, G. Wan, and Y . Chen, “Inference of brain states under anesthesia with meta learning based deep learning models,”IEEE Trans. Neural Syst. Rehabil. Eng., vol. 30, pp. 1081–1091, 2022
2022
-
[13]
Machine learning of EEG spectra classifies unconsciousness during GABAergic anesthesia,
J. H. Abel, M. A. Badgeley, B. Meschede-Krasaet al., “Machine learning of EEG spectra classifies unconsciousness during GABAergic anesthesia,”PLoS One, vol. 16, no. 5, p. e0246165, 2021
2021
-
[14]
Improved tracking of sevoflurane anesthetic states with drug-specific machine learning models,
K. Kashkooli, S. L. Polk, E. Y . Hahmet al., “Improved tracking of sevoflurane anesthetic states with drug-specific machine learning models,”J. Neural Eng., vol. 17, no. 4, p. 046020, 2020
2020
-
[15]
Multiclass brain- computer interface classification by Riemannian geometry,
A. Barachant, S. Bonnet, M. Congedo, and C. Jutten, “Multiclass brain- computer interface classification by Riemannian geometry,”IEEE Trans. Biomed. Eng., vol. 59, no. 4, pp. 920–928, 2012
2012
-
[16]
Riemannian approaches in brain- computer interfaces: A review,
F. Yger, M. Berar, and F. Lotte, “Riemannian approaches in brain- computer interfaces: A review,”IEEE Trans. Neural Syst. Rehabil. Eng., vol. 25, no. 10, pp. 1753–1762, 2017
2017
-
[17]
Transfer learning: A Riemannian geometry framework with applications to brain- computer interfaces,
P. Zanini, M. Congedo, C. Jutten, S. Said, and Y . Berthoumieu, “Transfer learning: A Riemannian geometry framework with applications to brain- computer interfaces,”IEEE Trans. Biomed. Eng., vol. 65, no. 5, pp. 1107–1116, 2018
2018
-
[18]
Transfer learning for brain-computer interfaces: A Euclidean space data alignment approach,
H. He and D. Wu, “Transfer learning for brain-computer interfaces: A Euclidean space data alignment approach,”IEEE Trans. Biomed. Eng., vol. 67, no. 2, pp. 399–410, 2020
2020
-
[19]
An introduction to ROC analysis,
T. Fawcett, “An introduction to ROC analysis,”Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006
2006
-
[20]
The use of the area under the ROC curve in the evaluation of machine learning algorithms,
A. P. Bradley, “The use of the area under the ROC curve in the evaluation of machine learning algorithms,”Pattern Recognit., vol. 30, no. 7, pp. 1145–1159, 1997
1997
-
[21]
Assessing the performance of prediction models: A framework for traditional and novel measures,
E. W. Steyerberg, A. J. Vickers, N. R. Cooket al., “Assessing the performance of prediction models: A framework for traditional and novel measures,”Epidemiology, vol. 21, no. 1, pp. 128–138, 2010
2010
-
[22]
Compara- tive effects of ketamine on bispectral index and spectral entropy of the electroencephalogram under sevoflurane anaesthesia,
P. Hans, P.-Y . Dewandre, J. F. Brichant, and V . Bonhomme, “Compara- tive effects of ketamine on bispectral index and spectral entropy of the electroencephalogram under sevoflurane anaesthesia,”Br . J. Anaesth., vol. 94, no. 3, pp. 336–340, 2005
2005
-
[23]
The effect of ketamine on clinical endpoints of hypnosis and EEG variables during propofol infusion,
T. Sakai, H. Singh, W. D. Mi, T. Kudo, and A. Matsuki, “The effect of ketamine on clinical endpoints of hypnosis and EEG variables during propofol infusion,”Acta Anaesthesiol. Scand., vol. 43, no. 2, pp. 212– 216, 1999
1999
-
[24]
Effects of ketamine and nitrous oxide on the bispectral index during propofol-fentanyl anaesthesia,
K. Hirota, T. Kubota, H. Ishihara, and A. Matsuki, “Effects of ketamine and nitrous oxide on the bispectral index during propofol-fentanyl anaesthesia,”Eur . J. Anaesthesiol., vol. 16, no. 11, pp. 779–783, 1999, pMID: 10713872
1999
-
[25]
Existence of multiple transitions of the critical state due to anesthetics,
D. Curic, D. M. Ashby, A. McGirr, and J. Davidsen, “Existence of multiple transitions of the critical state due to anesthetics,”Nat. Commun., vol. 15, p. 7025, 2024
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.