{"id":"ef88c8f8-eca4-4717-8f01-a42291b72752","arxiv_id":"2607.06629","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A JEPA-style EEG foundation model with shallow EMA targets plus light reconstruction reaches strong multi-task transfer and 3.06-year validation age MAE on a large multi-site corpus.","lead":"STST-JEPA is a self-supervised EEG transformer pretrained on ~48k sessions that predicts age with 3.06-year validation MAE from frozen embeddings and ranks first on a public multi-task EEG leaderboard. It is a practical foundation-model attempt for cheap brain-age and related biomarkers across pediatric-to-adult EEG.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Internal age MAE is best-validation on a pediatric-heavy split, and NeuralBench age rank-1 is window-length unmatched; foundation-model claim rests on these two soft metrics.","rationale":"The reader correctly isolates the load-bearing soft spot: the internal age result is best-validation on a pediatric-heavy combined corpus, and NeuralBench age rank-1 depends on unmatched 30 s windows. That is not a fabrication or an over-read; the manuscript itself states both caveats (Results; Methods §External Benchmark Protocol; Discussion). The multi-task NeuralBench evidence (especially sex and psychopathology remaining rank 1 at 2 s) and the large multi-site pretraining still support a foundation-model framing, so the verdict should stay CONDITIONAL rather than move to REJECT. No stronger internal inconsistency (e.g., leakage of the NeuralBench test into pretraining, or collapse of the joint objective) is visible from the text. The concrete fix the reader already wants—fixed internal test numbers plus matched-window primary leaderboard presentation—is exactly the check that would settle whether the age leg of the claim holds at the strength currently advertised. Agreement with the reader is therefore full; no verdict change is required beyond reinforcing CONDITIONAL.","tokens_in":20932,"tokens_out":733,"duration_ms":6964,"concrete_test":"Report a single fixed-protocol evaluation of the frozen attentive probe on the reserved internal subject-disjoint test partition (N=9,655 sessions), with separate brain.space-only and HBN-only MAEs/RMSEs/r, and re-rank the NeuralBench age entry under the official 2 s protocol as the primary number (30 s as secondary). If internal test MAE rises above ~4.5 yr combined or brain.space-only stays near 4.8+ yr and 2 s age remains rank 4, the foundation-model age claim should be demoted relative to the multi-task transfer claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that one pretrained encoder is a usable EEG foundation model, evidenced by (i) frozen-probe age MAE 3.06 yr (r=0.924) on 3,367 held-out validation sessions and (ii) rank-1 NeuralBench×brain.space results for sex/age/psychopathology with native 30 s windows. Both legs are soft in the same way the reader flags. Internal age is explicitly “best held-out validation … across the training trajectory” (Results §Validation Age Regression; Table 2), not a fixed internal test protocol; the internal test partition is reserved and only touched for exploratory BAG analysis. The combined MAE is also mechanically lowered by HBN’s pediatric mass (brain.space-only MAE is 4.82 yr). On NeuralBench, age rank-1 (r=0.749) uses unmatched 30 s windows; under the standard 2 s protocol the same checkpoint falls to rank 4 (r=0.691) while sex/psychopathology stay rank 1 (Results §External Benchmark Performance). The multi-task external evidence is real, but the headline age numbers that carry the “competitive age regression across the full pediatric-to-older-adult range” framing are the least protocol-matched pieces of the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces STST-JEPA, a 24-layer self-supervised EEG transformer pretrained on 47,703 sessions (ages 5–81) from brain.space and HBN. It combines a shallow-target latent prediction loss (predictor outputs matched to an EMA-of-tokenizer target under spatiotemporal block masks) with a down-weighted per-patch smooth-L1 reconstruction term, using coordinate-aware PMA channel pooling over a unified 128-channel budget. A frozen attentive probe reaches best held-out validation age MAE 3.06 yr (r=0.924) on 3,367 sessions; light final-layer fine-tuning yields rank-1 NeuralBench×brain.space placements for sex (bal. acc. 0.911), age (r=0.749), and psychopathology (r=0.215) with native 30 s windows. An exploratory bias-corrected brain-age-gap analysis shows small negative correlations with several cognitive-efficiency targets.","tokens_in":21291,"tokens_out":1602,"duration_ms":19750,"significance":"If the multi-task transfer results hold under matched protocols, this is a useful empirical contribution to EEG foundation models: one backbone, heterogeneous montages, pediatric-to-adult span, and public leaderboard rank-1 on three diverse labels. Strengths include subject-level stratified splits, detailed preprocessing and hyperparameter disclosure (Table 1), transparent caveats (pediatric-heavy MAE, 30 s vs 2 s windows, exploratory BAG, no clinical claim), and an external fixed-protocol benchmark. The shallow-target JEPA-plus-reconstruction design is a clear architectural choice relative to BENDR, LaBraM, EEG2Rep, and related work. The work is significant as an engineering and transfer study more than as a definitive brain-age biomarker paper.","major_comments":[{"comment":"Results §Validation Age Regression and Table 2: the headline internal age result is the best validation MAE across the training trajectory (3.06 yr combined; 4.82 yr brain.space-only), not a fixed internal test-protocol evaluation. The internal test partition is reserved and only used for exploratory BAG. For a foundation-model claim that leads with competitive age regression across 5–81 years, a single fixed-protocol subject-disjoint test number (or at least reporting the MAE at a pre-specified checkpoint) is load-bearing. Please report fixed-protocol test metrics for age on the reserved internal test set, or reframe the internal age result as a monitoring/stress-test metric rather than a primary endpoint.","section":"Results §Validation Age Regression; Table 2"},{"comment":"Results §External Benchmark Performance and Table 4: NeuralBench age rank-1 (r=0.749) uses unmatched 30 s windows against prior entries evaluated under the standard 2 s protocol; under 2 s the same checkpoint falls to rank 4 (r=0.691) while sex and psychopathology remain rank 1. The abstract and contributions list lead with rank-1 age without foregrounding this mismatch. The multi-task foundation claim is still supported by sex and psychopathology under matched 2 s windows, but age should not be presented as rank-1 without equal prominence of the 2 s result. Please restructure Table 4 / abstract so matched-protocol ranks are primary and 30 s results are secondary.","section":"Results §External Benchmark Performance; Table 4; Abstract"},{"comment":"Introduction and Methods §Self-Supervised Objective: the joint latent-plus-reconstruction objective is the paper’s central design choice, yet the manuscript states that controlled ablations were beyond compute budget and are not claimed empirically. Without at least a minimal ablation (latent-only vs reconstruction-only vs joint at fixed compute/steps) on a held-out probe, the claim that the joint formulation is the right inductive bias remains architectural intuition. A small-scale ablation on a subset of the corpus, or an honest demotion of the joint objective from a validated contribution to a design hypothesis, is needed for the Methods framing to match the evidence.","section":"Introduction; Methods §Self-Supervised Objective; Discussion"},{"comment":"Abstract and Introduction frame competitive age regression “across the full pediatric to older adult range” as a gap the work fills, but Results explicitly note that the combined 3.06 yr MAE is mechanically lowered by HBN’s pediatric mass and is not comparable to adult clinical benchmarks (e.g., TUAB 7–8 yr MAE). Brain.space-only MAE is 4.82 yr. The abstract should report brain.space-only (or age-stratified) MAE alongside the combined figure, and avoid implying adult-clinical comparability that the body correctly disclaims.","section":"Abstract; Introduction; Table 2"}],"minor_comments":[{"comment":"Figure numbering is inconsistent in the text: “Figure 1b” is described as the age distribution, then “Figure 2” and “Figure 3” captions appear to reuse “Figure 1b” / “Figure 2” labels in the manuscript dump. Please renumber figures and cross-references consistently.","section":"Results figures"},{"comment":"Table 0 is numbered before Table 1; consider renumbering cohort summary as Table 1 and shifting the hyperparameter table, or using a non-table summary box.","section":"Methods §Datasets"},{"comment":"Typographical variants of the model name (STST-JEP A / STST-JEPA) appear in the abstract and introduction; standardize.","section":"Abstract; Introduction"},{"comment":"BAG analysis (Table 5, Figure 3/4): effect sizes are small (|r|<0.10) and subject-level n is inflated; the paper already flags this. Consider moving the full 21-target panel to a supplement and keeping only FDR-surviving targets plus the n-inflation caveat in the main text.","section":"Results §Brain Age Gap"},{"comment":"Discussion projects falsifiable MAE ranges under 3× corpus scaling and fine-tuning; these are useful but should be clearly boxed as hypotheses so they are not read as secondary claims of the present work.","section":"Discussion"},{"comment":"Cite EEG2Rep more carefully when positioning shallow-target JEPA: the related-work paragraph already mentions it; a short sentence on how shallow EMA-tokenizer targets differ from EEG2Rep’s latent targets would help novelty placement.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The multi-task NeuralBench evidence under matched 2 s windows (sex and psychopathology rank 1) is the cleanest support for a foundation-model claim; the internal best-val age MAE and unmatched 30 s age rank-1 are the softest legs and currently over-weighted in the abstract. If the authors rebalance the abstract/Table 4 and add either a fixed internal test age number or a minimal objective ablation, this is a solid major-revision accept candidate for a methods/ML-for-neuro journal. Scope fit depends on whether the venue prioritizes foundation-model transfer or clinical biomarker validity—the manuscript correctly disclaims the latter."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: they pretrained a JEPA-style transformer on ~48k multi-site EEG sessions (brain.space + HBN, ages 5–81), then showed one backbone transfers to sex, age, and a psychopathology composite on a public leaderboard. That multi-task external evidence is the load-bearing part of the foundation-model claim, not the internal age number.\n\nWhat is actually new is the package more than any single trick: shallow EMA-of-tokenizer targets (not deep-encoder EMA), a down-weighted patch reconstruction term, and coordinate-aware PMA pooling so 115- and 128-channel montages share a 128-channel budget. Methods are careful—subject-level stratified splits, detailed ICA/ASR conditioning, hyperparameter tables, and unusually explicit caveats. Sex and psychopathology stay rank-1 even under the standard 2 s NeuralBench protocol; that is real transfer, not just longer windows.\n\nSoft spots match what the stress-test flags, and the authors mostly flag them too. The 3.06 yr MAE is best validation over the training trajectory, not a fixed internal test; the test set is reserved. Combined MAE is also helped by HBN’s pediatric mass (brain.space-only is 4.82 yr). NeuralBench age rank-1 uses unmatched 30 s windows and drops to rank 4 at 2 s. No joint-objective ablation, no public code/checkpoint, and part of the data is private. The BAG × cognition analysis is exploratory, small effects, directionally consistent—fine as secondary, not as clinical validation. They do not claim clinical validity.\n\nCitation pattern is appropriate for the EEG foundation-model and JEPA literature. Math is standard SSL, not a formal contribution. For people building EEG SSL or brain-age pipelines, this is worth reading; for general ML, it is a domain application paper.\n\nI would send it to peer review. Ask for fixed-protocol internal test numbers, primary reporting of matched-window leaderboard results, and ideally code/checkpoints plus at least a light objective ablation. Engage if you work in this area; skip if you only care about architectural novelty.","headline":"Solid multi-task EEG foundation-model paper; the multi-site pretraining and NeuralBench transfer are real, while the 3.06 yr age headline is softer than the abstract sells.","tokens_in":21988,"tokens_out":554,"would_cite":true,"duration_ms":10541,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"One self-supervised EEG transformer pretrained on 47,703 sessions becomes a foundation model that predicts age to 3.06 years MAE and ranks first on three public downstream tasks from a shared backbone.","keywords":["EEG","self-supervised learning","brain age","foundation model","JEPA","transformer","spatiotemporal masking","montage heterogeneity"],"falsifier":"Retrain or re-evaluate under a single fixed subject-disjoint internal test protocol and under the leaderboard’s standard 2-second windows for all three tasks; if age MAE and rank collapse while sex/psychopathology do not, the foundation claim rests mainly on evaluation choices and window length rather than the joint objective and backbone.","tokens_in":21744,"feed_emoji":"🧠","tokens_out":1025,"duration_ms":20261,"temperature":0.7,"pith_summary":"This paper argues that EEG can support a practical brain-age biomarker if a single pretrained encoder can absorb montage differences, limited labels, and subject-level non-stationarity across childhood through older adulthood. STST-JEPA is a 24-layer transformer trained with a joint objective: predict latent representations of masked spatiotemporal patches against a shallow EMA tokenizer target, while a down-weighted reconstruction term keeps those latents faithful to the waveform. On frozen embeddings, a lightweight attentive probe reaches a best held-out validation age MAE of 3.06 years (r = 0.924) on 3,367 sessions, far below a predict-the-mean baseline near 10 years. With light final-layer fine-tuning, the same checkpoint places rank 1 on the public NeuralBench × brain.space leaderboard for sex, age, and a psychopathology composite using the model’s native 30-second windows. The age residual, after bias correction, also trends negatively with cognitive efficiency across several tasks, tying the representation to behavioral performance beyond chronological age alone.","feed_headline":"One EEG backbone predicts age to 3.06 years MAE","feed_subtitle":"Same pretrained encoder ranks first on sex, age, and psychopathology from 30-second windows.","key_machinery":"Shallow-target spatio-temporal JEPA: masked-token latent prediction against an EMA-of-tokenizer target (not a deep-encoder twin), plus a lighter per-patch signal reconstruction loss, with coordinate-aware pooled multihead attention that collapses arbitrary channel montages into a fixed 128-channel token grid.","core_discovery":"A single STST-JEPA encoder pretrained without age labels on 47,703 multi-site EEG sessions can serve as an EEG foundation model: frozen embeddings support competitive age regression across ages 5–81, and light task-specific fine-tuning of final layers yields rank-1 leaderboard performance on sex classification, age prediction, and psychopathology regression from one shared backbone.","pith_inferences":["If corpus scaling and light age-aware pretraining auxiliaries deliver the projected gains the authors sketch, EEG brain-age models may become practical screening tools long before clinical biomarker claims are warranted.","The same shallow-target-plus-reconstruction recipe may transfer to other sparse, montage-heterogeneous biosignals (e.g., multi-site MEG or wearable ECG) where pure reconstruction overfits artifact energy.","A resting-only BAG analysis would cleanly separate trait age structure from task-evoked dynamics that currently mix into the residual–efficiency correlations.","Rank-1 multitask transfer from one checkpoint is stronger evidence of foundation status than any single age MAE, so future work should prioritize fixed multitask protocols over chasing absolute age error alone."],"forward_implications":["One frozen EEG backbone can supply competitive probes for age, sex, and paradigm identity without retraining the encoder.","Light final-layer fine-tuning of that backbone can transfer to multiple public leaderboard tasks, including a long-horizon psychopathology composite.","Bias-corrected brain-age residuals can carry small but directionally consistent links to speeded cognitive efficiency beyond chronological age.","Coordinate-aware channel pooling plus a unified channel budget can absorb cross-site montage mismatch without forcing a single physical electrode layout.","A deliberately down-weighted reconstruction term can act as a soft floor that keeps latent prediction from drifting away from the waveform."],"fun_headline_variants":["STST-JEPA frozen embeddings hit 3.06-year MAE on EEG age","Pretrained STST-JEPA ranks first on sex, age, psychopathology","One EEG encoder regresses age at 3.06 MAE across ages 5-81","Shared STST-JEPA backbone tops NeuralBench from 30s windows","Self-supervised STST-JEPA yields r=0.924 age prediction frozen"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the headline age result fairly proves foundation-model quality when it is the best validation score over training—not a fixed internal test protocol—on a pediatric-heavy corpus whose narrower age spread lowers MAE, while the leaderboard age win uses longer 30-second windows than the standard 2-second protocol.","fun_headline_variants_meta":{"raw":{"variants":["STST-JEPA frozen embeddings hit 3.06-year MAE on EEG age","Pretrained STST-JEPA ranks first on sex, age, psychopathology","One EEG encoder regresses age at 3.06 MAE across ages 5-81","Shared STST-JEPA backbone tops NeuralBench from 30s windows","Self-supervised STST-JEPA yields r=0.924 age prediction frozen"]},"model":"grok-4.5","effort":"low","cost_usd":0.00638,"raw_usage":{"total_tokens":1724,"prompt_tokens":895,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":63800000,"prompt_tokens_details":{"text_tokens":895,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":737,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":895,"tokens_out":92,"duration_ms":6426,"temperature":1.0,"reasoning_tokens":737,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T01:06:11.502394+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain or re-evaluate under a single fixed subject-disjoint internal test protocol and under the leaderboard’s standard 2-second windows for all three tasks; if age MAE and rank collapse while sex/psychopathology do not, the foundation claim rests mainly on evaluation choices and window length rather than the joint objective and backbone.","supporting_citations":[],"review_version":1}