{"id":"b76e3d5c-86b8-41bb-a11c-696c2f47d229","arxiv_id":"2605.13930","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Sparse autoencoders on EEG transformers extract clinical features, identify three steering regimes, expose age-pathology entanglements and wrecking-ball failures, and map interventions to frequency spectra.","lead":"The paper applies TopK Sparse Autoencoders to three EEG foundation models to extract and clinically ground sparse features from their embeddings. A smart generalist might read it to learn how to open up black-box medical AI for brain signals and spot when those models mix up important clinical concepts.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Clinical taxonomy may be incomplete, making observed entanglements artifacts of unmeasured confounders rather than model failures","rationale":"The reader’s weakest assumption directly identifies the load-bearing point for the central claim about representational failures. Because the full manuscript is now available, the same assumption remains the least secure link; no other internal inconsistency (e.g., in SAE training or spectral decoder) appears more critical on a first read.","tokens_in":1727,"tokens_out":363,"duration_ms":14646,"concrete_test":"Re-run the concept-steering experiments after augmenting the taxonomy with two additional variables (sleep stage and recording site) extracted from the same datasets; recompute the target vs. off-target probe area metric and the age-pathology entanglement score. If either metric changes by >15% or the number of “selectively steerable” features increases, the original taxonomy was insufficient.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (wrecking-ball interventions and age-pathology confounding where suppression of one corrupts the other) requires that the four-concept taxonomy supplies a sufficient, unbiased basis for monosemanticity measurement. If the underlying EEG datasets contain correlations between the labeled concepts and unmeasured variables (e.g., recording site, specific sleep-stage distributions, or unannotated comorbidities), then the “encoded but entangled” regime and the steering selectivity metric could arise from taxonomy incompleteness rather than intrinsic representational structure. The single intrinsic dictionary-health hyperparameter procedure is asserted to transfer across SleepFM, REVE, and LaBraM, but this assumes the health audit itself is architecture-agnostic and that TopK SAE features remain stable under that choice; neither is shown to be robust to alternative health metrics or to a richer taxonomy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper applies TopK Sparse Autoencoders to the embeddings of three EEG foundation models (SleepFM, REVE, LaBraM) to extract sparse feature dictionaries. These features are grounded in a four-concept clinical taxonomy (abnormality, age, sex, medication) to measure monosemanticity and entanglement. A single intrinsic dictionary-health hyperparameter procedure is used across architectures. Concept steering is performed with a new 'target vs. off-target' probe area metric that identifies three regimes (selectively steerable, encoded but entangled, non-encoded). The work reports 'wrecking-ball' interventions that collapse performance and specific clinical entanglements (e.g., age-pathology confounding), with a spectral decoder translating interventions into frequency-domain signatures.","tokens_in":1894,"tokens_out":624,"duration_ms":30186,"significance":"If the empirical results and metric definitions hold under the stated taxonomy, the framework supplies a concrete, transferable auditing procedure for representational quality in clinical EEG models and directly links latent interventions to physiologically interpretable spectral changes. The architecture-agnostic hyperparameter procedure and the steering selectivity metric are potentially reusable contributions.","major_comments":[{"comment":"The central claims of 'encoded but entangled' regimes and age-pathology confounding (abstract and the steering results section) rest on the four-concept taxonomy supplying a sufficient, unbiased basis for monosemanticity measurement. The manuscript does not report controls for unmeasured confounders (recording site, sleep-stage distributions, or comorbidities) that could induce the observed steering non-selectivity; without such checks the entanglement findings risk being artifacts of taxonomy incompleteness rather than intrinsic model structure.","section":"clinical taxonomy grounding and steering results"},{"comment":"The claim that a single intrinsic dictionary-health hyperparameter procedure 'transfers robustly across all three architectures' (abstract and methods) is load-bearing for the cross-model generality result. No ablation is shown against alternative health metrics or an expanded taxonomy, leaving open whether TopK feature stability is an artifact of the chosen audit rather than a general property.","section":"hyperparameter procedure and cross-architecture results"}],"minor_comments":[{"comment":"The definition and computation of the 'target vs. off-target probe area metric' should be given explicitly with a formula or pseudocode, including how the area is normalized and how statistical significance is assessed.","section":"steering selectivity metric"},{"comment":"Figure captions and axis labels for the spectral decoder outputs should explicitly state the frequency bands corresponding to 'pathological slow-wave suppression' and 'α-band restoration' so readers can map them to standard EEG conventions without ambiguity.","section":"spectral decoder figures"}],"recommendation":"major_revision","confidential_remarks":"The citation pattern for prior SAE work in vision/language is appropriate, but the manuscript should clarify whether any of the three EEG models were developed by the authors; if so, that relationship should be disclosed for scope fit with the journal."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which highlight important considerations for the robustness of our taxonomy and hyperparameter choices. We respond to each major comment below.","responses":[{"response":"We agree that the four-concept taxonomy does not include explicit controls for unmeasured confounders such as recording site, sleep-stage distributions, or comorbidities, and the manuscript does not report such checks. The available dataset annotations are limited to the four concepts, so additional controls would require new data or metadata not present in the public releases. The observed entanglements are consistent across three independent models, which supports that they reflect representational properties, but we acknowledge this does not fully rule out dataset artifacts. We will add a limitations subsection discussing the taxonomy scope and the potential influence of unmeasured confounders on steering selectivity.","revision_made":"yes","referee_comment":"The central claims of 'encoded but entangled' regimes and age-pathology confounding (abstract and the steering results section) rest on the four-concept taxonomy supplying a sufficient, unbiased basis for monosemanticity measurement. The manuscript does not report controls for unmeasured confounders (recording site, sleep-stage distributions, or comorbidities) that could induce the observed steering non-selectivity; without such checks the entanglement findings risk being artifacts of taxonomy incompleteness rather than intrinsic model structure."},{"response":"The dictionary-health metric is intrinsic to the SAE optimization and does not depend on the clinical taxonomy labels, which is the basis for claiming transfer without per-model retuning. We demonstrate this empirically on three architecturally distinct models. We did not include ablations against alternative health metrics or an expanded taxonomy. We will revise the methods section to elaborate on the metric's selection rationale from prior SAE literature and to note the absence of such ablations as a limitation and direction for future work.","revision_made":"partial","referee_comment":"The claim that a single intrinsic dictionary-health hyperparameter procedure 'transfers robustly across all three architectures' (abstract and methods) is load-bearing for the cross-model generality result. No ablation is shown against alternative health metrics or an expanded taxonomy, leaving open whether TopK feature stability is an artifact of the chosen audit rather than a general property."}],"tokens_in":1445,"tokens_out":476,"duration_ms":40988,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper takes TopK sparse autoencoders and runs them on the embeddings of SleepFM, REVE, and LaBraM. It grounds the resulting features in four clinical labels, measures how cleanly each feature aligns with one label versus the others, and tests steering by seeing how much changing one feature affects the target concept without hitting the others. A spectral decoder then turns those changes into frequency-band effects. That combination of SAE application, the selectivity metric, and the spectral readout is new for EEG foundation models. The pipeline is straightforward and the regimes they describe (selectively steerable, encoded but entangled, non-encoded) give a usable way to talk about what the models actually represent. The single hyperparameter procedure that works across architectures is also a practical plus if it holds up. The main soft spot is the clinical taxonomy itself. If age, pathology, sex, and medication are correlated with unmeasured factors such as recording site, sleep stage distribution, or comorbidities, then the reported entanglements and wrecking-ball effects could be artifacts of incomplete labeling rather than intrinsic model structure. The abstract does not show ablations on richer taxonomies or alternative health metrics, so it is hard to tell how robust the findings are. The work is worth a serious referee if the full manuscript contains the ablation tables, error bars, and dataset details needed to check those points. People building or deploying EEG models would get value from seeing whether the steering results replicate.","headline":"They apply TopK SAEs to three EEG transformers, add a target-vs-off-target steering metric, and map interventions to spectra, but the entanglement claims depend on a narrow clinical taxonomy that may miss confounders.","tokens_in":2443,"tokens_out":376,"would_cite":false,"duration_ms":18524,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"EEG SAE interpretability pipeline (monosemanticity taxonomy, concept steering, wrecking-ball regimes) has no structural overlap with RS J-cost, φ-ladder or 8-tick forcing","alignment":"orthogonal","rationale":"The paper's core machinery (TopK SAEs on EEG embeddings, TCAV attribution, excess-selectivity metric Δ̃, spectral decoder) is standard mechanistic-interpretability tooling applied to clinical EEG concepts. It never invokes the reciprocal cost J(x)=½(x+x⁻¹)−1, golden-ratio identities, recognition-cost forcing, 8-tick periodicity, or any theorem from the RS chain (e.g., reality_from_one_distinction, washburn_uniqueness_aczel, alexander_duality_circle_linking). Domain (EEG transformers) and methodology are disjoint from RS; no contradiction arises either.","tokens_in":48907,"confidence":"high","tokens_out":192,"duration_ms":6583,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Sparse autoencoders reveal entangled clinical concepts in EEG foundation models, such as age and pathology confounding.","keywords":["EEG foundation models","sparse autoencoders","mechanistic interpretability","concept steering","monosemanticity","clinical entanglement","age-pathology confounding","spectral decoder"],"falsifier":"A steering intervention on a pathology-labeled feature that alters age predictions in a manner inconsistent with the measured entanglement level would falsify the claim of quantifiable clinical entanglements.","tokens_in":2654,"feed_emoji":"🧠","tokens_out":590,"duration_ms":32044,"temperature":0.7,"pith_summary":"The paper uses TopK sparse autoencoders on embeddings from three different EEG foundation models to learn sparse dictionaries of features. These features are then evaluated against a clinical taxonomy of abnormality, age, sex, and medication to measure how cleanly each concept is represented. A single hyperparameter selection method based on dictionary health works across all models. Concept steering experiments identify features that can be changed selectively, those that are entangled, and those not present, while also showing interventions that destroy overall model performance. A decoder translates the feature changes into changes in brain wave frequency spectra.","feed_headline":"Sparse autoencoders expose entangled concepts in EEG models","feed_subtitle":"Steering experiments show some features can be changed in isolation while age and pathology cannot be separated without side effects.","key_machinery":"TopK Sparse Autoencoders that produce sparse feature dictionaries from model embeddings, paired with a target vs. off-target probe area metric for measuring steering selectivity.","core_discovery":"TopK SAEs extract features from EEG transformer embeddings that can be grounded in clinical concepts, revealing three regimes of encoding and exposing failures where concept steering either collapses global performance or entangles concepts like age and pathology such that one cannot be altered without the other.","pith_inferences":["Models may require additional training to disentangle clinical variables before deployment in targeted interventions.","The framework could help identify which features are safe to manipulate in clinical settings.","Extending the spectral mapping might allow direct prediction of EEG changes from model edits."],"forward_implications":["Some concepts allow selective steering without off-target effects.","Other concepts are encoded but entangled, preventing isolated intervention.","Certain interventions act as wrecking balls that collapse model performance globally.","The spectral decoder maps latent feature changes to interpretable amplitude spectrum shifts like slow-wave suppression.","Clinical entanglements such as age-pathology confounding make independent suppression impossible."],"fun_headline_variants":["SAEs expose three encoding regimes in EEG models","Clinical entanglements block isolated steering in EEG transformers","TopK SAEs ground features to age and pathology in EEG","Steering failures collapse EEG model performance globally"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The clinical taxonomy of abnormality, age, sex, and medication is sufficient and unbiased for measuring how monosemantic the extracted features are.","fun_headline_variants_meta":{"raw":{"variants":["SAEs expose three encoding regimes in EEG models","Clinical entanglements block isolated steering in EEG transformers","TopK SAEs ground features to age and pathology in EEG","Steering failures collapse EEG model performance globally"]},"model":"grok-4.3","cost_usd":0.003784,"raw_usage":{"total_tokens":1940,"prompt_tokens":638,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":37837000,"prompt_tokens_details":{"text_tokens":638,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1242,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":638,"tokens_out":60,"duration_ms":9989,"temperature":1.0,"reasoning_tokens":1242,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T06:03:07.221486+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A steering intervention on a pathology-labeled feature that alters age predictions in a manner inconsistent with the measured entanglement level would falsify the claim of quantifiable clinical entanglements.","supporting_citations":[],"review_version":3}