{"id":"b97912f9-9312-40fe-8071-7d71bf012ce2","arxiv_id":"2607.12364","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Pixel and CLIP metrics fail for EEG image reconstructions; a VLM-distilled BCI-Coherence Score better tracks perceptual-semantic recoverability and human agreement.","lead":"Standard image metrics misjudge EEG-to-image reconstructions that are blurry but still meaningful. This work proposes BCS, a VLM-based score that separates perceptual fidelity from recoverable meaning and tracks human judgments more closely.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Circularity risk: BCS is distilled from four VLMs' T-PAS/T-SAS consensus and then scored against those same targets, so reported MAE/r may largely measure self-prediction rather than independent validity.","rationale":"The reader's weakest_assumption correctly isolates the load-bearing premise: VLM consensus is both the supervisory target and the scoring reference for BCS. Abstract-only access leaves no evidence of held-out splits, leave-one-VLM-out distillation, or primary human-vs-BCS metrics, so the circularity concern stands and keeps the verdict CONDITIONAL with low confidence. No stronger independent concern is available from the abstract; the human agreement numbers are promising but secondary. The concrete hold-out test would settle whether the headline MAE/r survive without self-prediction.","tokens_in":2194,"tokens_out":436,"duration_ms":3564,"concrete_test":"Hold out a random 20% of the 6855 pairs never seen by the BCS distillation; recompute MAE and Pearson r of BCS against T-PAS/T-SAS on that hold-out only. If MAE rises above ~0.15 or r falls below ~0.5, the reported figures are largely self-prediction and the claim weakens. Separately report BCS-human correlation on the same hold-out.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on BCS achieving T-PAS MAE 0.079 (r=0.700) and T-SAS MAE 0.082 (r=0.850). Because BCS is distilled from the four VLMs' consensus that also defines T-PAS/T-SAS, those numbers can be inflated by training-target leakage unless the distillation uses strictly held-out pairs, models, or questions. Human kappa/alpha=0.882 is the only external signal and is only summarized; without non-circular splits or primary human-vs-BCS correlation, the MAE/r do not yet establish that BCS is a valid independent evaluator of perceptual-semantic recoverability for degraded EEG reconstructions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript argues that standard pixel and representation metrics (SSIM, LPIPS, CLIP) are poorly suited to EEG-to-image reconstruction, which is typically blurry and low-detail: they either over-penalize semantically recoverable outputs or reward plausible but wrong ones. On 6,855 ground-truth/reconstruction pairs from ATM, ENIGMA, BrainVis, and DreamDiffusion, the authors report near-zero correlation between pixel metrics and semantic consistency, and show that representation metrics conflate perceptual and semantic errors. They propose a BCI-aware evaluation framework in which four VLMs answer structured questions to produce Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS); the VLM consensus is distilled into a compact BCI-Coherence Score (BCS) that achieves T-PAS MAE 0.079 (r=0.700) and T-SAS MAE 0.082 (r=0.850). Human joint-coherence agreement is reported as Cohen's kappa = 0.882 ± 0.174 and Krippendorff's alpha = 0.082, supporting perceptual-semantic recoverability as the appropriate evaluation target.","tokens_in":2371,"tokens_out":1032,"duration_ms":7363,"significance":"If the circularity and validation concerns can be resolved, the work would supply a practically useful, BCI-specific evaluator for a setting where generic vision metrics systematically misrank reconstructions. The multi-method analysis (semantic probes, caption harshness/blind-spot rates, controlled degradations) on a large multi-model pair set is a genuine contribution to evaluation methodology for neural decoding. Distilling VLM consensus into a compact BCS and releasing code/resources would lower the barrier to adoption. The human reliability numbers, if they reflect independent human-vs-BCS agreement rather than only inter-annotator agreement on the joint task, would be strong external support. The central idea—that EEG reconstruction evaluation should prioritize recoverable meaning over pixel fidelity—is well motivated and field-relevant.","major_comments":[{"comment":"The central quantitative claim for BCS (T-PAS MAE 0.079, r=0.700; T-SAS MAE 0.082, r=0.850) is at high risk of circularity. BCS is distilled from the four VLMs' consensus that also defines T-PAS/T-SAS, then scored against those same targets. Without strictly held-out pairs, held-out VLMs, held-out question templates, or primary evaluation against independent human labels (not VLM-derived scores), the reported MAE/r can largely measure self-prediction of the supervisory signal rather than independent validity of BCS as an evaluator of perceptual-semantic recoverability. This must be clarified with explicit train/eval splits and non-circular metrics before the claim can be accepted.","section":null},{"comment":"Human validation is summarized only as joint-coherence agreement (kappa = 0.882 ± 0.174, alpha = 0.882). It is not clear whether this is inter-annotator reliability on the joint task, human agreement with BCS, or human agreement with T-PAS/T-SAS. For BCS to be established as a valid independent evaluator, the manuscript needs primary human-vs-BCS correlations (and preferably human-vs-T-PAS/T-SAS) on a held-out set of degraded reconstructions, with protocol details (number of raters, sampling of pairs, rating scale, blinding). Inter-annotator reliability alone does not validate the automated score.","section":null},{"comment":"The premise that four VLMs answering structured questions yield valid T-PAS/T-SAS labels for heavily degraded EEG reconstructions is load-bearing and currently under-supported in the abstract. VLM judgments on blurry, distorted, low-detail images can be unstable or systematically biased. The manuscript should report (i) inter-VLM agreement, (ii) failure cases / blind-spot rates of the VLMs themselves on this domain, and (iii) sensitivity of T-PAS/T-SAS and BCS to ensemble membership and question templates. Without that, the supervisory target itself remains an untested axiom.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":"Abstract-only review: methods, splits, prompts, architecture, and failure cases are unavailable, so confidence is low. The circularity concern is structural and will not be fixed by presentation alone; if the full paper already has held-out VLMs/pairs and primary human-vs-BCS results, the recommendation can move to minor_revision. If BCS is only validated against its own distillation targets, the central claim does not hold and reject would be appropriate. Scope is a reasonable fit for a CV/BCI evaluation venue if the validation is cleaned up."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is a methods paper arguing that SSIM/LPIPS/CLIP are the wrong tools for blurry EEG reconstructions, and that we should score tolerant perceptual vs semantic recoverability instead. On the abstract alone that diagnosis is plausible and field-relevant.\n\nWhat looks new is the framing and the artifacts. They run a large-ish failure analysis (6,855 pairs from ATM, ENIGMA, BrainVis, DreamDiffusion) showing near-zero pixel–semantic correlation and representation metrics that mix the two error types. They then define T-PAS and T-SAS via structured VLM questions, distill a compact BCS from four-VLM consensus, and report human joint-coherence agreement (kappa/alpha ~0.88). If the full paper ships prompts, code, and splits as the site claims, that is real reusable work for BCI evaluation, not just another CLIP wrapper.\n\nSoft spots, in proportion. The stress-test concern is real on the abstract: BCS is distilled from the same VLM consensus it is scored against (MAE 0.079 / r 0.700 for T-PAS; 0.082 / 0.850 for T-SAS). Without held-out VLMs, held-out pairs, or primary human-vs-BCS correlation, those numbers can partly be self-prediction. Human agreement is the stronger external signal, but it is only summarized here—no protocol, no failure cases, no BCS–human table. Free parameters (which four VLMs, question templates, distillation setup) are also invisible. So the central claim is not yet independently established; it is conditional on the full methods.\n\nI would not treat the MAE/r as settled validity. I would treat the problem statement and the human-agreement direction as worth a serious look. This is for people who train or evaluate EEG/BCI image reconstruction and care about ranking models honestly. It deserves a serious referee if the full paper shows non-circular splits and human primary validation; desk-reject only if those are missing. I would bring it to reading group as a methods discussion, not as a settled result. Cite only after checking the validation design.","headline":"Useful evaluation proposal for EEG-to-image, but abstract-only and circularity risk on BCS MAE/r keep confidence low.","tokens_in":3077,"tokens_out":536,"would_cite":false,"duration_ms":4337,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"EEG-to-image reconstruction should be scored by recoverable meaning, not pixel similarity; a VLM-consensus BCI-Coherence Score tracks that better.","keywords":["EEG-to-image reconstruction","perceptual-semantic coherence","BCI evaluation","vision-language models","BCI-Coherence Score","T-PAS","T-SAS","semantic recoverability"],"falsifier":"A new set of EEG reconstructions whose human-rated joint coherence diverges sharply from BCS (or from the VLM T-PAS/T-SAS targets), or an ablation showing BCS MAE/r collapses once the same VLMs are removed from the distillation loop.","tokens_in":3029,"feed_emoji":"🧠","tokens_out":569,"duration_ms":4503,"temperature":0.7,"pith_summary":"EEG-to-image pipelines produce blurry, distorted outputs that still sometimes preserve the original meaning. Standard metrics such as SSIM, LPIPS, and CLIP either punish those recoverable images for looking imperfect or reward images that look plausible but mean the wrong thing. The authors examine thousands of ground-truth and reconstruction pairs from four recent systems and show that pure pixel metrics barely track semantic consistency while representation metrics mix visual and meaning errors together. They therefore build a BCI-aware evaluation loop in which four vision-language models answer structured questions about each pair, yielding tolerant perceptual and semantic scores whose consensus is distilled into a single BCI-Coherence Score (BCS). BCS closely matches the VLM targets and aligns with human joint-coherence judgments, arguing that the right success criterion for brain-to-image work is perceptual-semantic recoverability rather than generic visual likeness.","feed_headline":"EEG image scores track meaning, not pixels","feed_subtitle":"A VLM-distilled BCI-Coherence Score matches human recoverability judgments far better than SSIM or CLIP.","key_machinery":"The BCI-Coherence Score (BCS), obtained by distilling the consensus of four VLMs that answer structured questions to produce Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS) for each ground-truth/reconstruction pair.","core_discovery":"Pixel-level and representation-level metrics fail to separate visual fidelity from recoverable meaning on EEG reconstructions; a compact BCI-Coherence Score distilled from four VLMs' structured Tolerant Perceptual Alignment Scores and Tolerant Semantic Alignment Scores recovers those two dimensions with low error and high human agreement, establishing recoverability as the proper evaluation target.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Pixel metrics fail EEG images while VLMs track recoverable meaning","BCS from four VLMs matches human EEG reconstruction judgments","Separate fidelity from semantics for EEG-to-image scoring","SSIM and CLIP misrank EEG reconstructions; BCS does not","Recoverability not similarity is the right EEG image metric"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That four vision-language models answering structured questions give valid perceptual and semantic labels for heavily degraded EEG reconstructions, and that scoring the distilled BCS against those same labels is not largely circular self-prediction.","fun_headline_variants_meta":{"raw":{"variants":["Pixel metrics fail EEG images while VLMs track recoverable meaning","BCS from four VLMs matches human EEG reconstruction judgments","Separate fidelity from semantics for EEG-to-image scoring","SSIM and CLIP misrank EEG reconstructions; BCS does not","Recoverability not similarity is the right EEG image metric"]},"model":"grok-4.5","effort":"low","cost_usd":0.008412,"raw_usage":{"total_tokens":1946,"prompt_tokens":816,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":84120000,"prompt_tokens_details":{"text_tokens":816,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1047,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":816,"tokens_out":83,"duration_ms":7756,"temperature":1.0,"reasoning_tokens":1047,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T06:40:43.139747+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A new set of EEG reconstructions whose human-rated joint coherence diverges sharply from BCS (or from the VLM T-PAS/T-SAS targets), or an ablation showing BCS MAE/r collapses once the same VLMs are removed from the distillation loop.","supporting_citations":[],"review_version":1}