{"id":"c81966f4-d573-4852-a459-67ccdf59d7ef","arxiv_id":"2607.06630","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Lipschitz-style certificates pass for EEGNet and classical BCI decoders while PGD accuracy drops up to 25.7% at ε=0.25, and related fidelity and privacy audits reveal further objective-user welfare mismatches.","lead":"Robustness certificates for EEG brain-computer interface models can pass while classification accuracy falls by about 26% under small adversarial noise. The work shows that formal math checks alone do not guarantee safe real-world neural interfaces and offers a three-part audit for that gap.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The verification-insufficiency claim rests on a certificate the paper itself shows is vacuous for task safety, so the gap is partly by construction rather than a surprising failure of formal methods.","rationale":"The reader correctly isolates the representativeness of the global product-of-layer certificate as the weakest assumption. The empirical accuracy drops under PGD and physiological stress are solid and architecture-independent for task fragility; the soft spot is whether those drops demonstrate a failure of formal verification as it would actually be used for safety claims, or only of a deliberately vacuous sufficient condition. The paper’s own margin-certified-safe fraction of zero and §6.2 admission of conservativeness make this internal rather than external. A local-Lipschitz re-audit is the single check that settles it. No stronger objection (e.g., data leakage or statistical invalidity of E1) is needed; E2/E3 are secondary. Verdict remains CONDITIONAL: useful empirical audit, but the headline governance claim should be tempered until non-vacuous certificates are shown to exhibit the same gap.","tokens_in":16978,"tokens_out":610,"duration_ms":6718,"concrete_test":"Re-run the official-session EEGNet E1 grid (§4.1, §5.2) with a tighter local Lipschitz estimator (e.g., CLEVER-style or SDP/local bounds around each clean epoch) and recompute both bare lemma pass rate and margin-certified-safe fraction at ε=0.25. If the local certificate still passes for all 9 subjects while PGD accuracy drop remains ~0.257, the insufficiency claim strengthens; if the local certificate fails whenever task accuracy collapses, the gap is largely an artifact of the global product-of-layer bound.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central E1 claim (abstract; §5.2; Table 1; Fig. 2) is that a Lipschitz-style certificate remains valid while official-session accuracy drops 25.7% at ε=0.25 (and similarly for CSP/FBCSP). Lemma 1 (§3) only bounds output movement by Lε; the paper then reports that the margin-certified-safe fraction is zero for every nonzero budget (§4.1, §5.2) and later acknowledges the product-of-layer bound is known to be conservative (§6.2). Thus the certificate that “passes” is the bare output-sensitivity check, which the authors already treat as insufficient for task correctness. Classical CSP/FBCSP results show task fragility but are not Lipschitz-certified models, so they do not independently establish that a non-vacuous formal certificate can pass while the task fails. If the only certificates that pass are those already known to be too loose to certify decisions, the verification-insufficiency result is weaker than the deployment claim that “operational safety auditing, not certificate verification alone, is necessary.”","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that formal robustness certificates for embedded neural-interface models can be mathematically valid while task performance collapses under bounded perturbations, and treats this as one instance of a broader alignment failure between training proxies and user welfare. It defines three operational failure modes—verification insufficiency (Definition 1), proxy-fidelity divergence (Definition 2), and latent information exfiltration (Definition 3)—and instantiates them as empirical audits E1–E3 on BCI Competition IV 2a (official train-to-evaluation sessions) and SEED-IV (session-holdout). The load-bearing E1 result is that at ε=0.25, official-session EEGNet accuracy drops by 0.2573 under PGD while a conservative Lipschitz-style output-sensitivity check remains passed for all 9 subjects; parallel accuracy drops are reported for CSP+LDA and FBCSP+LDA. E2 shows objective-dependent fidelity trade-offs (time-domain auxiliary loss improves MSE by 0.1132 while worsening spectral log-MSE), and E3 shows public-task embeddings leak subject identity at ~48% versus ~6.7% chance, with permutation controls collapsing to chance. The authors conclude that operational multi-objective auditing, not certificate verification alone, is necessary for responsible deployment.","tokens_in":17337,"tokens_out":1492,"duration_ms":15524,"significance":"If the central claim holds, the paper supplies a concrete, falsifiable demonstration that a standard sufficient-condition certificate can be operationally uninformative for EEG decoding, and packages that demonstration with classical baselines, official-session protocols, null controls, and paired statistical tests. The architecture-independent task-failure pattern (EEGNet, CSP+LDA, FBCSP+LDA) and the multi-metric E2/E3 audits are useful engineering contributions for BCI safety practice, even if the alignment framing is more rhetorical than theoretical. Strengths include the official MOABB session protocol, label- and private-attribute permutation controls, subject-paired sign-flip tests with BH correction, and explicit reporting that the margin-certified-safe fraction is zero. The work is significant as a methodological audit paper rather than as a new certification method or a general theory of alignment.","major_comments":[{"comment":"The load-bearing E1 claim (abstract; §5.2; Table 1; Fig. 2) rests on a certificate the paper itself shows is vacuous for task safety. Lemma 1 (§3) only bounds output movement by Lε; §4.1 and §5.2 report that the margin-certified-safe fraction is zero for every nonzero budget, and §6.2 acknowledges the product-of-layer bound is known to be conservative. The “certificate that passes” is therefore the bare output-sensitivity check, which the authors already treat as insufficient for decision correctness. Classical CSP/FBCSP results establish task fragility under PGD but are not Lipschitz-certified models, so they do not independently show that a non-vacuous formal certificate can pass while the task fails. The deployment conclusion that “operational safety auditing, not certificate verification alone, is necessary” overstates what the current certificate comparison supports. Either (i) repl","section":null},{"comment":"The architecture-independence claim for the verification gap (§1 contributions; abstract; §5.2) is imprecise. The paper correctly shows that PGD task failure is architecture-independent, but the formal certificate C(fθ,ε) is only computed for the spectral-normalized / EEGNet setting (§4.1). CSP+LDA and FBCSP+LDA are used as task-failure controls, not as certified models. The abstract and contribution statements should be revised so that “verification gap” is not attributed to the classical pipelines, and so that the architecture-independent claim is limited to adversarial task degradation.","section":null},{"comment":"E3 statistical support is weaker than the abstract’s 48.1% vs 6.7% claim suggests (§4.4; §5.4; Table 3; Table 5). With n=3 seeds the minimum exact one-sided p-value is 0.125; the paper correctly labels this as directional multiseed evidence, but the abstract and conclusion present the leakage figure without that caveat. RF/MLP probes are capacity-bounded and class-balanced subsampled (n_train_fit=6000), so measured leakage may understate what a fully optimized probe could extract. Either increase seeds / report full-data nonlinear probes, or qualify the abstract claim to match the inferential limits already stated in §4.4 and §6.2.","section":null}],"minor_comments":[{"comment":"Figure 1 is schematic and does not convey quantitative results; consider moving it to an appendix or tightening the caption so it does not compete with Figs. 2–5 for attention.","section":null},{"comment":"Notation for the certificate is slightly inconsistent: C(fθ,ε) in Definition 1 vs. the bare lemma check in Eq. (8) and the √2 Lε margin rule in §4.1. A short clarifying sentence would help.","section":null},{"comment":"The SEED-IV E3 experiments use pre-extracted de LDS features rather than raw EEG (§5.1, §6.2). This is disclosed but should also appear in the abstract or early methods so readers do not assume raw-signal privacy leakage.","section":null},{"comment":"Table 4 and Fig. 5 are useful but dense; highlighting the physiological conditions that meet or exceed the ε=0.25 PGD drop (already bolded in the table note) in the main text would improve readability.","section":null},{"comment":"A few phrasing issues: “deceptive alignment” is used as an operational analogy (§2.1, §3.2) for non-strategic EEG decoders; a single clarifying sentence that this is analogy, not a claim of mesa-optimization, would reduce misreading.","section":null},{"comment":"References [8], [9], [14], [20] are alignment surveys/position pieces; their use is fine for framing, but the empirical sections should continue to stand without them.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical audit paper dressed in alignment language. The E1 result is real and well controlled for what it is (conservative global certificates + PGD task failure), but the current framing invites the skeptic’s objection that the gap is partly by construction. If the authors tighten the certificate or narrow the claim, this is a useful BCI/ML-safety contribution; if they insist on the strong “certificates fail / formal verification alone is insufficient” headline without a non-vacuous certificate comparison, the paper will continue to draw the same critique. Scope fit is reasonable for a methods/safety venue that accepts empirical BCI work; less so for a pure verification venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: on official BCI IV 2a train-to-eval sessions, EEGNet accuracy falls about 25.7% at ε=0.25 under PGD while their product-of-layer Lipschitz output-sensitivity check still passes for all nine subjects, and CSP+LDA / FBCSP+LDA show the same task fragility. That empirical package is the paper’s real contribution, not a new theory of certificates.\n\nWhat is actually new is the quantified architecture-independent gap under the official protocol, plus the joint packaging of three audits (verification, proxy-fidelity, privacy) with null controls and paired stats. E1 is the load-bearing piece and it is done carefully: session-level splits, label-permutation near chance, BH-corrected subject-paired tests, classical baselines. E2’s objective-dependent trade-offs (time-domain MSE improves while spectral log-MSE worsens, and the reverse under spectral losses) are clean and controlled. E3 shows subject identity above chance with a permutation collapse; under-powered, but directionally honest.\n\nThe soft spot the stress-test flags is real and not minor: Lemma 1 only bounds output movement, the margin-certified-safe fraction is zero for every nonzero budget, and the authors themselves call the bound conservative. So the certificate that “passes” is already the one they treat as insufficient for task correctness. Classical pipelines show fragility but are not Lipschitz-certified, so they do not prove that a non-vacuous formal certificate can pass while the task fails. The deployment slogan (“operational auditing, not certificates alone”) is still reasonable as engineering advice; it is just stronger than the formal gap they actually measured. Scope is narrow (two EEG datasets, de LDS features for privacy, no released code), and the alignment vocabulary is heavier than the experiments require.\n\nMath is standard sufficient-condition Lipschitz; data and citations look solid for the claims as written. This is for people writing BCI safety cases and for neurotech/AI-safety readers who need multi-objective audits, not for pure verification theorists. I would send it to peer review. Engage if you work on BCI robustness or neuroprivacy; treat the certificate story as a demonstration of loose bounds plus task fragility, not as a surprise failure of formal methods.","headline":"Clean official-session evidence that task accuracy collapses under PGD while a deliberately loose Lipschitz check still passes; useful audit package, but the verification gap is partly by construction.","tokens_in":17915,"tokens_out":570,"would_cite":true,"duration_ms":16136,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Formal robustness certificates for neural-interface models can pass while task accuracy collapses under attack, so operational safety audits—not certificates alone—are required for responsible deployment.","keywords":["neural interfaces","robustness certificates","EEG decoding","AI alignment","privacy leakage","proxy fidelity","BCI safety","operational auditing"],"falsifier":"Re-run the official-session EEGNet PGD grid with a tighter local Lipschitz or margin certificate: if that certificate yields a nonzero certified-safe fraction while accuracy at ε=0.25 stays near clean levels—or fails whenever accuracy collapses—the claimed operational gap between certification and task safety is closed.","tokens_in":17859,"feed_emoji":"🧠","tokens_out":731,"duration_ms":17836,"temperature":0.7,"pith_summary":"This paper argues that mathematical safety certificates for brain–computer interface models can look valid while the models fail at the job users care about: decoding intent accurately under stress. At a moderate attack budget, EEGNet motor-imagery accuracy falls by about a quarter while a Lipschitz-style output-sensitivity certificate still passes for every subject tested, and the same task-failure pattern appears in classical CSP and FBCSP pipelines. The authors treat this gap as one instance of a broader alignment failure—training proxies diverging from user welfare—and propose a unified empirical audit around three linked modes: certificates that pass while behavior degrades, task-trained representations that damage neural signal structure, and public-task embeddings that leak private subject identity. Instantiated on standard EEG benchmarks with official session splits, null controls, and paired tests, the results show that responsible deployment needs multi-objective operational auditing, not certificate verification alone.","feed_headline":"Brain-decoder certificates pass while accuracy falls 25.7%","feed_subtitle":"Operational safety audits, not math certificates alone, are required for neural interfaces.","key_machinery":"A unified empirical audit framework organized around three failure modes: verification insufficiency (certificate true while task-level safety false), proxy-fidelity divergence (task objective improves while a separate neural-fidelity metric degrades), and latent information exfiltration (public-task embeddings retain recoverable private attributes). The framework is carried by official session-level protocols, null controls, and paired statistical tests on EEG/BCI data.","core_discovery":"At perturbation budget ε=0.25, official-session EEGNet classification accuracy drops by 25.7% under projected-gradient attack while a Lipschitz-style certificate remains valid for all nine BCI Competition IV 2a subjects; the same task-failure pattern holds for CSP+LDA and FBCSP+LDA, so the verification gap is architecture-independent. Parallel audits show that a time-domain auxiliary objective can cut reconstruction MSE by 0.1132 while worsening spectral log-MSE, and that public-task embeddings recover subject identity at 48.1% versus 6.7% chance. Operational safety auditing across verification, fidelity, and privacy—not certificate verification alone—is therefore necessary for responsible n","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Certificates hold as EEGNet accuracy drops 25.7% under attack","Architecture-independent gap: certificates pass while accuracy falls","Verification gap across EEGNet CSP FBCSP as accuracy collapses 25.7%","Certificates valid for all 9 subjects yet task accuracy falls 25.7%","Operational safety audits needed beyond certificates alone"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The central claim rests on treating a loose global product-of-layer Lipschitz upper bound as a representative formal certificate that would be offered or accepted as safety evidence; if tighter certificates closed the gap, the verification-insufficiency result would weaken.","fun_headline_variants_meta":{"raw":{"variants":["Certificates hold as EEGNet accuracy drops 25.7% under attack","Architecture-independent gap: certificates pass while accuracy falls","Verification gap across EEGNet CSP FBCSP as accuracy collapses 25.7%","Certificates valid for all 9 subjects yet task accuracy falls 25.7%","Operational safety audits needed beyond certificates alone"]},"model":"grok-4.5","effort":"low","cost_usd":0.005076,"raw_usage":{"total_tokens":1486,"prompt_tokens":864,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":50760000,"prompt_tokens_details":{"text_tokens":864,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":532,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":864,"tokens_out":90,"duration_ms":5982,"temperature":1.0,"reasoning_tokens":532,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T01:03:09.334002+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the official-session EEGNet PGD grid with a tighter local Lipschitz or margin certificate: if that certificate yields a nonzero certified-safe fraction while accuracy at ε=0.25 stays near clean levels—or fails whenever accuracy collapses—the claimed operational gap between certification and task safety is closed.","supporting_citations":[],"review_version":1}