{"id":"eb127e43-cd07-4e44-8c45-4f7355eb4187","arxiv_id":"2607.22778","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 165 subjects and 216,714 pipeline evaluations, no single EEG decoding pipeline wins for everyone; compact portfolios of 12 pipelines recover 90-96% of the per-subject oracle.","lead":"Researchers benchmarked 216,714 EEG motor-imagery decoding pipelines across three public datasets and found that the best pipeline varies sharply from person to person. They show that small portfolios of pipelines can recover most of the per-subject best performance, but only if the right pipeline can be selected for each new user.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The heterogeneity and portfolio-gain evidence is not benchmarked against a null model; with ~1000 noisy pipeline scores per subject, distinct-argmax counts and max-of-K oracle retention may reflect sampling noise rather than true subject-specific pipeline complementarity.","rationale":"The reader correctly identified that oracle-in-set presupposes a selector and therefore represents an upper bound on achievable personalization. That is a valid practical caveat, but the more load-bearing issue is internal to the evidence: the observed 'heterogeneity' and the portfolio gains may arise entirely from the statistical mechanics of selecting maxima from a large set of noisy scores. The paper counts 42/93 distinct winners, but with ~1000 pipelines per subject, distinct argmaxes are expected even when pipelines are equivalent; the count has no stated null distribution. Similarly, oracle-retention is the mean per-subject max over the portfolio; adding pipelines to the portfolio raises the expected max under pure noise. The K=1 baseline is also chosen using test labels, so the headline 'single best global pipeline retains 94.2%' is itself an oracle-selected quantity. Without a permutation or noise model, the central claim that subject-level heterogeneity is real and exploitable is unverified. This is a distinct concern from the one the reader emphasized: even if a perfect selector existed, the portfolio numbers may simply reflect noise maximization. I therefore recommend withholding a definitive verdict until the null-model check is performed. The benchmark itself is large, standardized, and internally consistent; the family-level descriptive differences (cov-tgsp/CSP, secondary metrics, Friedman results) are real and valuable. But those findings do not establish that the subject-specific optimal pipeline is stable or that compact portfolios exploit genuine complementarity.","tokens_in":26283,"tokens_out":7393,"duration_ms":87040,"concrete_test":"Run the same 10×80/20 ShuffleSplit pipeline on a null matrix where, for each pipeline, subject scores are permuted across subjects (preserving each pipeline's marginal score distribution and breaking subject-specific structure). Recompute (a) the number of distinct winning pipelines per fold and (b) Top-K Mean oracle-retention for K=1 and K=12, for each dataset-band. If the observed 42/93 distinct winners and the 96.5%/90.0% retentions fall within the permutation-null 95% interval, the heterogeneity/exploitability claim is not supported; if they fall outside, the claim survives this confound.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"Section 3.2 reports 42/93 distinct winning pipelines, and Section 3.5/Figure 4 reports oracle-retention gains (94.2→96.5% Cho2017; 81.8→90.0% PhysionetMI) as evidence that heterogeneity is exploitable. Both statistics come from subject-by-pipeline matrices with ~864–1048 scores per subject. Under the null that every pipeline has the same expected accuracy for a given subject, the argmax is essentially random; with that many near-tied noisy scores, nearly every subject will have a different 'winner' even without any subject-specific structure. Likewise, oracle-in-set is the mean over test subjects of the max of K noisy per-subject scores, so it mechanically increases with K and approaches the global oracle, independent of any true subject×pipeline interaction. The K=1 baseline itself is selected on test data (Section 2.6: 'fixed best global pipeline was defined as the single pipeline with the highest mean test performance'), making it an oracle-selected reference rather than a train-selected fixed pipeline. The paper reports no permutation or synthetic-null control for either the distinct-winner counts or the retention curves. Consequently, the central claim that the landscape is subject-dependent and 'can be exploited' is not yet distinguished from selection noise. The missing selector acknowledged in Section 4 is a practical limitation; the absence of a null model is an evidential one.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale, standardized within-session benchmark of EEG motor imagery decoding pipelines across three public datasets (Cho2017: 52 subjects, PhysionetMI: 109 subjects, Zhou2016: 4 subjects), using the MOABB LeftRightImagery framework, two frequency bands, six feature families, multiple preprocessing steps, and classical/MLP classifiers. The authors report that cov-tgsp and CSP are the strongest families overall, but that the identity of the best full pipeline varies strongly across subjects (42 distinct winners across 52 Cho2017 subjects; 93 across 109 PhysionetMI subjects). They then construct compact portfolios of pipelines from the benchmark and report oracle-retention ratios, claiming a single best global pipeline retains 94.2%/81.8% of the subject-specific oracle and a K=12 Top-K Mean portfolio raises this to 96.5%/90.0%. The paper concludes that the decoding landscape is subject-dependent and that this heterogeneity can be exploited through compact portfolios.","tokens_in":26696,"tokens_out":4717,"duration_ms":53719,"significance":"If the heterogeneity and portfolio claims are supported, the paper would be a useful contribution to BCI benchmarking: it provides a large, standardized, publicly reproducible evaluation of a broad pipeline space, checks robustness across accuracy/balanced-accuracy/AUC/F1/precision/recall, and proposes a concrete way to reduce a large search space to a compact set of candidate pipelines. The availability of code and the use of held-out subjects in portfolio evaluation are clear strengths. However, the central interpretive claims about “exploitable” subject-level heterogeneity currently rest on oracle-based quantities and lack any null-model comparison, so the practical conclusion is not yet established.","major_comments":[{"comment":"The evidence for “true” subject-level heterogeneity and portfolio complementarity is not benchmarked against a null model. The subject-by-pipeline matrices contain roughly 864–1048 scores per subject. Under the null that each pipeline has the same expected accuracy for a given subject (no subject×pipeline interaction), the argmax is essentially random: with that many near-tied noisy scores, nearly every subject will have a different “winner”, and the max-of-K oracle-in-set mechanically increases with K. The reported distinct-winner counts (42/52, 93/109) and retention gains (94.2→96.5% Cho2017; 81.8→90.0% PhysionetMI) are therefore not sufficient to establish that the landscape is subject-dependent or that the gains reflect real complementarity. Please add a permutation or synthetic-null control (e.g., permuting subject labels across pipelines, or simulating scores from subject and pipel","section":"§3.2, Table 2; §3.5, Fig. 4"},{"comment":"The K=1 baseline is not a feasible fixed-pipeline baseline. Section 2.6 defines the “fixed best global pipeline” as the pipeline with the highest mean test performance, i.e., it is selected after seeing the held-out test subjects. Consequently, the reported retention at K=1 (94.2%/81.8%) is an oracle-selected upper bound, not the performance of a train-selected single pipeline. In addition, the portfolio quantities oracle-in-set and oracle-retention ratio (Eq. 1) are upper bounds: no selector exists to choose the best pipeline for a new subject, as Section 4 concedes. The conclusions in Section 5, which state that heterogeneity is “practically exploitable” through compact portfolios, overstate what the current analysis supports. Please supply a train-selected K=1 baseline and/or explicitly frame all portfolio results as upper bounds pending a calibration-based or transfer-based selector.","section":"§2.6, §3.5, §5"},{"comment":"The family-level statistical comparisons use, for each subject, the best-performing pipeline within each family, but the number of pipeline variants differs substantially across families and bands (e.g., Cov+TGSP 8–30 Hz in Cho2017 has only 16 pipelines, while most other family-band cells have 88). Best-of-family scores systematically favor families with more variants, so the Friedman/Wilcoxon comparisons in §3.4 do not test family quality per se. This is a load-bearing issue for the descriptive claim that cov-tgsp and CSP are the “strongest” families. Please either control for the number of variants (e.g., matched subsampling or per-family model-selection estimates) or explicitly state and justify the best-of-family aggregation, or report the sensitivity of the family ranking to this choice.","section":"§3.4, Table 1"}],"minor_comments":[{"comment":"The abstract states 44,928, 109,000, and 4,192 subject-level observations, while the full-text abstract states 61,464, 132,762, and 22,488. The latter are raw-row counts; please reconcile the terminology to avoid confusion.","section":"Abstract vs. §3.1"},{"comment":"Several free parameters are stated without justification, including HFD maximum scale k=10, the post-cue window 0.6–2.0 s, and the two frequency bands. A short sensitivity analysis or citation for these choices would strengthen the paper.","section":"§2.2.2, §2.4"},{"comment":"The right-column panels have an axis label that appears cut off (“ balanced accuracy”), and the y-axis in panels A/C/E should be explicitly labeled as the oracle-retention ratio. Also, the shaded regions are fold-wise standard deviations; a note on how many folds are included would help.","section":"Figure 4"},{"comment":"The “Ties” column reports tie cases before tie-breaking, but the tie-breaking rule is not described. Please state how ties were resolved for the distinct-winner counts.","section":"Table 2"},{"comment":"The portfolio-selection strategies (greedy, regret_greedy, submodular coverage, etc.) are described verbally; the composite objective is not given a precise formula. Please include the exact objective used for the non-Top-K strategies, or point to the code with a stable version/DOI, so the results are reproducible.","section":"§2.6"},{"comment":"The manuscript uses “PhysionetMI” and “PhysionetMotorImagery” interchangeably; please standardize. There are also duplicated references (e.g., [50] and [64] both list Cho2017; [57] repeats [15]).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The descriptive benchmark is likely a solid and useful contribution, and the code availability is a plus. The main risk is interpretive: the portfolio claims are currently supported only by oracle-based quantities and lack a null model. There is no indication of fabrication or circularity; the issues are missing controls and an overstating of the practical claim. A revision that adds a null model, a train-selected baseline, and toned-down conclusions could make the paper acceptable. I would not recommend rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is a serious large-scale benchmark that's worth citing for the resource, but the claim that the heterogeneity is exploitable is not yet supported without a null model.\n\nWhat's actually new: the subject-level winning-pipeline counts and the portfolio-retention curves. Those aren't in the MOABB references. The portfolio framing is a good idea, and the robustness check across seven selection strategies shows the result isn't an artifact of one heuristic. The finding that useful diversity lives mostly within the dominant family rather than across families is interesting and runs against the initial intuition. The code and data are public, and the paper is honest about the within-session protocol and about not having a selector for new users.\n\nWhere it's soft. The biggest gap is the missing null model. With roughly a thousand pipeline scores per subject, the argmax pipeline is extremely sensitive to noise. Under a null where all pipelines have the same expected accuracy, you'd still see nearly every subject with a different 'best' pipeline. The same applies to the oracle-in-set curves: max-of-K over noisy per-subject scores rises with K mechanically. The paper reports no permutation or synthetic-null control, so the heterogeneity and retention gains are not yet separated from selection noise.\n\nSecond, the K=1 baseline is selected on the test data—defined as the pipeline with highest mean test performance. That makes the baseline optimistic, so the gains over K=1 are conservative, not generous. But the absolute retention ratios are upper bounds because they assume an oracle within the portfolio. The paper itself concedes there's no mechanism to choose among the portfolio for a new user. These two biases pull in opposite directions and the paper doesn't untangle them.\n\nThird, the family-level significance tests compare the best pipeline within each family, with unequal family sizes. That's a best-of-N comparison, so the pairwise tests are not clean.\n\nAlso, the abstract and Section 3.1 give different aggregated observation counts (216,714 vs 158,120). Minor, but it should be fixed.\n\nNone of this sinks the descriptive benchmark. The family rankings—cov-tgsp and CSP on top—look robust across datasets and metrics. The 'exploitability' conclusions, though, need work before I'd trust them.\n\nWho this is for: anyone working with MOABB or building pipeline-selection resources will want the benchmark and code. The portfolio analysis is useful as a case study in oracle-based evaluation. I'd give it a serious referee, and ask for a permutation null, a train-selected baseline, and a clear statement that the portfolio gains are upper bounds. With those changes it could be a solid contribution.\n\nSend it to review.","headline":"The benchmark is a valuable reusable resource; the exploitability claim needs a null model and a train-selected baseline before it's credible.","tokens_in":27111,"tokens_out":5717,"would_cite":true,"duration_ms":59772,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across hundreds of subjects, the best EEG motor-imagery decoding pipeline changes per person; a compact portfolio of twelve recovers most of the per-subject best.","keywords":["EEG","motor imagery","brain-computer interface","pipeline benchmarking","subject heterogeneity","portfolio selection","Common Spatial Patterns","Riemannian tangent space"],"falsifier":"Test the portfolio on new subjects with a practical selection rule, such as a two-minute calibration block evaluated on the K=12 pipelines, and compare realized accuracy against the oracle-in-set value; if the realized gain over the single best pipeline is zero or negative, the exploitable-heterogeneity claim is refuted. Alternatively, compare Top-K Mean against a random K=12 portfolio: if random selection matches its retention, the gain is not attributable to subject-level structure.","tokens_in":26204,"feed_emoji":"🧠","tokens_out":4159,"duration_ms":47711,"temperature":0.7,"pith_summary":"The paper argues that EEG motor-imagery decoding has no universal best pipeline: aggregate rankings conceal that the optimal full configuration changes markedly from subject to subject. Across 52 subjects in one dataset and 109 in another, 42 and 93 distinct winning pipelines were observed, and the best feature family varied as well. The authors then use the exhaustive benchmark as an empirical landscape to build compact portfolios of pipelines. A single best global pipeline retains 94.2% of the per-subject oracle in the 52-subject dataset and 81.8% in the 109-subject dataset; a portfolio of just 12 pipelines raises retention to 96.5% and 90.0%, respectively. This matters because it suggests that personalization can be approached by choosing among a small set of strong candidate pipelines rather than by exhaustive search or a one-size-fits-all decoder.","feed_headline":"Portfolio of 12 pipelines beats the single best EEG pipeline","feed_subtitle":"Across 161 subjects the winning pipeline is subject-specific; 12 pipelines retain 90–96% of the per-subject best.","key_machinery":"The central object is the subject-by-pipeline performance matrix, built from 216,714 raw benchmark evaluations across three datasets, two frequency bands, six feature-extraction families (covariance tangent-space projection, CSP, coherence-based tangent space, Hjorth, Higuchi fractal dimension, SVD entropy), multiple scalers, and several classifiers. Portfolio construction uses repeated 80/20 subject-level splits: the Top-K Mean heuristic selects the K pipelines with the highest mean balanced accuracy on training subjects, and the portfolio is evaluated on held-out subjects via the oracle-in-set metric, defined as the mean across test subjects of the best score achieved by any pipeline in th","core_discovery":"The decoding landscape is subject-dependent, and this heterogeneity is practically exploitable. The paper shows that on Cho2017 (52 subjects) 42 different full pipelines were each the best for some subject, and on PhysionetMI (109 subjects) 93 were; even the winning feature family was not stable. Using the benchmark itself as a performance landscape, the authors construct portfolios: the Top-K Mean heuristic, which selects the K pipelines with the highest mean training balanced accuracy, retains 96.5% of the subject-specific oracle on Cho2017 and 90.0% on PhysionetMI at K=12, compared with 94.2% and 81.8% for the single best global pipeline. The diversity that drives these gains lies mostly","pith_inferences":["The paper defines the oracle-in-set but provides no mechanism to choose the best pipeline for a new subject in real time; the reported retention ratios are therefore upper bounds, not achieved personalization, until a practical selector is built.","The 21 ties observed on PhysionetMI before tie-breaking suggest that top pipelines are often statistically interchangeable; a random portfolio of the same size could capture a substantial part of the gain, which would weaken the claim that the landscape has exploitable subject-level structure.","Because the selected portfolios are dominated by cov-tgsp variants, a cheaper extension would be to test whether hyperparameter diversity alone within that one family saturates the oracle-retention curve.","A concrete next step is to use a short calibration block per new subject to pick among the portfolio members, then compare realized accuracy with the oracle-in-set; this would directly test the practical value of the portfolio."],"forward_implications":["A single fixed global pipeline leaves a measurable, dataset-dependent gap relative to the per-subject best; the gap is largest in the most heterogeneous dataset.","Reducing the search space to about twelve pipelines retains 90-96.5% of the subject-specific oracle, making personalization computationally feasible.","The useful diversity in portfolios comes mainly from within the dominant feature family, so practitioners can focus tuning effort on scalers and classifiers inside that family.","The results reframe so-called BCI illiteracy: poor decoding performance may reflect a pipeline-subject mismatch rather than an inherent inability of the user.","The natural next test is whether the same portfolio logic holds under cross-session and participant-independent evaluation protocols."],"fun_headline_variants":["EEG decoding's dirty secret: best pipeline varies by person","12 pipelines outperform any single EEG decoder across subjects","Portfolio of 12 EEG pipelines captures 96% of per-person best","Subject-specific EEG decoding: a portfolio beats the best single","One EEG pipeline never wins: 42 winners for 52 subjects"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The results assume a selector can identify, for each new subject, the best-performing pipeline inside the portfolio; the paper states it does not yet provide such a mechanism, so without it the portfolio's 90-96.5% retention is an upper bound rather than achieved personalization.","fun_headline_variants_meta":{"raw":{"variants":["EEG decoding's dirty secret: best pipeline varies by person","12 pipelines outperform any single EEG decoder across subjects","Portfolio of 12 EEG pipelines captures 96% of per-person best","Subject-specific EEG decoding: a portfolio beats the best single","One EEG pipeline never wins: 42 winners for 52 subjects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1311,"prompt_tokens":928,"completion_tokens":383,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":311}},"tokens_in":672,"tokens_out":383,"duration_ms":4935,"temperature":1.0,"reasoning_tokens":311,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:45:27.508078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Test the portfolio on new subjects with a practical selection rule, such as a two-minute calibration block evaluated on the K=12 pipelines, and compare realized accuracy against the oracle-in-set value; if the realized gain over the single best pipeline is zero or negative, the exploitable-heterogeneity claim is refuted. Alternatively, compare Top-K Mean against a random K=12 portfolio: if random selection matches its retention, the gain is not attributable to subject-level structure.","supporting_citations":[],"review_version":1}