{"id":"b4324f53-32e6-4a4f-a1c6-17bb165bebca","arxiv_id":"2607.27857","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"EEG-image retrieval models that score high on standard 200-way ranking often fail to pick the viewed image over controlled edits that change one visual factor, with attribute changes (shape, color, texture) hardest for all eight models tested.","lead":"This paper introduces EEG-EditBench, a benchmark that tests whether EEG-to-image retrieval models can still pick the image a person viewed when the alternatives are near-copies altered in one specific way — object swapped, color changed, background replaced, or object removed. On eight published models, strong performance at standard 200-way retrieval does not transfer to these controlled-edit tests, and fine-grained attribute changes are hardest.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Edit-family hierarchy is unverified once FLUX.2 regeneration artifacts are controlled: original-over-null 2AFC reaches 49.8–78.1% (Table S12), yet the null-vs-semantic control is reported only overall, not per family.","rationale":"We agree with the reader's weakest assumption and make it more precise. The central contribution is not merely the observation that 200-way Top-1 exceeds EP-Top1; it is the factor-level diagnosis that attribute edits are hardest and identity/removal easiest. For that diagnosis to hold, 2AFC differences between families must be attributable to the changed semantic factor. The paper's own null-edit data (Table S12) establishes that models can detect FLUX.2 regeneration artifacts (original-over-null up to 78.1%), so the source-vs-edit metric conflates semantic change with generation artifacts. The null-vs-semantic comparison controls for this because both candidates are generated, but it is only reported overall. A per-family version is a natural, cheap, and decisive check. If attribute edits remain hardest when both candidates are generated, the hierarchy survives; if not, the benchmark's main claim must be re-scoped to 'edits with larger visual/artifact shifts are easier' rather than 'identity is better represented than attributes.' This does not undermine the more modest claim that standard retrieval does not transfer to edit-based evaluation, so the conditional verdict stands. The paper deserves credit for running null and distance-balancing controls, sharing code/data, and reporting subject-level CIs; our concern is about interpretation, not measurement integrity.","tokens_in":20340,"tokens_out":6215,"duration_ms":62507,"concrete_test":"Recompute the Table S12 protocol as a per-family statistic: for each of the four edit families, measure null-over-semantic 2AFC (source-concept EEG scoring the FLUX.2 null-edited image above each semantic edit) using the same null edits and semantic edits, with subject-level bootstrap CIs. Then compare the family ordering against Table 1's original-vs-edit 2AFC. If Attribute is not the lowest family in null-over-semantic correct-EEG 2AFC for a majority of the eight models, the 'attribute hardest' claim is not established; if it remains lowest (and pairing gains per family stay positive), the artifact concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's central factor-level finding — identity/removal edits are easier than attribute edits (Table 1; §4) — is load-bearing for the paper's claim that EEG-EditBench reveals which visual information models preserve. That inference requires that 2AFC gaps across families reflect the targeted visual factor rather than the image editor's artifacts or visual edit magnitude. The paper's own controls are incomplete. Figure 5 shows attribute/background edits have the smallest OpenCLIP shifts while identity/removal have the largest, so the family hierarchy co-varies with edit magnitude; Table S14 distance-balancing reduces but does not eliminate the gaps. More directly, Table S12 shows original-over-null 2AFC of 49.8–78.1% across models: eight systems can separate the original from a no-op FLUX.2 regeneration, so ordinary source-vs-edit 2AFC is contaminated by a detectable 'original vs generated' signal. The paper does report a generation-artifact control — null-over-semantic 2AFC with both candidates generated (Table S12) — but only as a micro-average over all 2,137 edits. It is not broken down by edit family, so the main hierarchy (identity/removal > background > attribute) has not been shown to survive when both candidates pass through the editing pipeline. Without that per-family breakdown, the headline 'attribute changes present the greatest challenge' could be an editor artifact (or a visual-magnitude effect) rather than a property of EEG-image representations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EEG-EditBench, a benchmark for probing which visual information EEG-to-image retrieval models actually use. Starting from the 200 THINGS-EEG2 test images, the authors use InternVL3.5 plus FLUX.2 to generate 2,137 human-reviewed controlled edits in four families (object identity, attribute, background, object removal) and evaluate eight recent EEG decoding models with their native similarity functions. The central claim is that high standard 200-way retrieval accuracy does not imply the ability to distinguish the viewed image from its controlled edits, and that across models object identity and removal edits are easier than fine-grained attribute edits. The benchmark also reports source-category effects and target-concept alignment, and supplies code and data.","tokens_in":20529,"tokens_out":2838,"duration_ms":32656,"significance":"If the central claim holds, this is a useful and timely benchmark: it identifies a blind spot in standard EEG-to-image retrieval evaluation and offers a public, reproducible tool for measuring it. The manuscript is strong in execution in several ways: eight independently reproduced models are evaluated under a common protocol; the edit-generation pipeline is clearly documented; human quality control is used; and the supplementary material includes multiple validity controls (mismatched-EEG pairing gains, null-edit comparisons, distance balancing, aggregation sensitivity). These controls are a real strength and make the benchmark substantially more credible than a simple collection of generated edits. However, the factor-level interpretation, especially the headline 'attribute changes are hardest' result, is not yet established to the standard the paper aims for because the generation-artifact control is not broken down by edit family.","major_comments":[{"comment":"The load-bearing factor-level claim in the Abstract and §4 ('fine-grained attribute changes presenting the greatest challenge') requires that the 2AFC gap between edit families reflects the intended visual factor rather than editor artifacts. The paper's own null-edit control shows that models can separate originals from no-op FLUX.2 regenerations by 49.8–78.1% (Table S12), and the null-vs-semantic control is reported only as an overall micro-average over all 2,137 edits, not per family. If the regeneration-artifact signal is stronger for identity/removal edits than for attribute edits, the observed hierarchy would be an artifact of the editing pipeline. Please report the null-over-semantic 2AFC (and ideally the original-over-null 2AFC) separately for identity, attribute, background, and removal edits, with confidence intervals, and state whether the family ranking survives when both can","section":"§S5.2, Table S12"},{"comment":"Distance balancing addresses visual edit magnitude but does not establish factor-level validity. Figure 5 shows that attribute and background edits have the smallest OpenCLIP cosine shifts while identity and removal edits have the largest, so the family hierarchy co-varies with edit magnitude. Table S14 shows that distance balancing reduces but does not eliminate the identity/removal-vs-attribute gaps. However, visual distance is not equivalent to artifact prevalence or to the specific 'original vs generated' signal detected in Table S12. The fact that the attribute-vs-background gap is close to zero for ATM after balancing, while the identity-vs-attribute gap remains positive in most models, leaves room for artifact-driven explanations. A per-family null-vs-semantic comparison, or a distance- and artifact-matched subset analysis, is needed before the 'attribute changes are hardest' conc","section":"Fig. 5 and §S6.1, Table S14"},{"comment":"The human QC excludes obvious artifacts and disqualifies edits that change unrelated content, but the large original-over-null 2AFC values (up to 78.1% for Brain-HIVE) show that models detect subtle regeneration differences even in human-approved no-op edits. This is not a flaw by itself, but it means the benchmark's 'valid edited image' status is not sufficient to guarantee that an edit changes only the target factor. The paper should state explicitly what the accepted non-target variation is, and should provide per-family statistics on visual distance, null-edit sensitivity, and any artifact-related metadata (e.g., reviewer flags). This would let readers assess whether the family hierarchy is robust to the editor's uncontrolled variation.","section":"§S2.3 and Table S12"}],"minor_comments":[{"comment":"The text says 'attribute changes are consistently more difficult than changes in object identity or presence' but does not mention that background edits sit between them in most models and that the ordering between background and attribute is not uniform (e.g., ATM). Please soften or qualify the sentence to match Table 1.","section":"§5 (Discussion)"},{"comment":"The figure uses a dashed line for the cross-concept reference and shaded IQR, but the boxplots are not described in the caption (whisker definition, outliers). Adding this information would improve reproducibility of the plot's interpretation.","section":"Fig. 5"},{"comment":"The column headers are split across two lines and the dagger/plus symbols are not defined. Please define 'Background−Attribute' notation in the caption and clarify the units.","section":"Table S14"},{"comment":"The text says Brain-HIVE 'visual caches are isolated by seed because its VAE branch includes stochastic sampling.' It would be helpful to state whether the SynCLR and CLIP branches are deterministic and cached across seeds, since this affects the meaning of seed variance.","section":"§S3.2"},{"comment":"Some references have future-venue placeholders (e.g., Jo et al. 2026, Zheng et al. 2026) that may not be finalized by the time of publication; please update them. Also, the dataset URL in the abstract appears as a GitHub repository; include a versioned DOI or archive identifier for long-term accessibility.","section":"Global"},{"comment":"The phrase 'controlled edits' is used consistently, but the paper does not quantitatively verify factor purity beyond human review and CLIP distance. A short sentence in §3 acknowledging the limitations of human review for subtle attribute edits would help set expectations.","section":"§1 and Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed benchmark paper with a meaningful message, but the central factor-level conclusion is not yet fully supported. The key missing piece is a per-family breakdown of the generation-artifact control (Table S12). If the authors can show that identity/removal vs attribute differences survive when both candidates are generated images and are matched for plausible artifact presence, the paper would be a solid accept. As it stands, the gap between the strength of the claim and the strength of the evidence merits a major revision rather than rejection, because the omission is fixable within the current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this paper ships a genuinely new diagnostic: 2,137 human-reviewed image edits across four families built on THINGS-EEG2, and it shows that standard 200-way Top-1 overstates what EEG-image retrieval models can actually distinguish. Second, the headline hierarchy — attribute edits are hardest — is not yet artifact-controlled, so treat that as provisional.\n\nThe dataset construction is careful: InternVL3.5 proposes edits, FLUX.2 generates, two independent human reviewers plus a third adjudicator approve each one. Eight models are reproduced with their native encoders and frozen visual features, five seeds per subject, and results are aggregated sensibly. The pairing-gain control (correct vs mismatched EEG, +13.9 to +35.1 pp) is strong evidence that scores aren't just image-prior noise. The null-edit control is the right instinct, and the distance-balanced comparison shows the identity/removal vs attribute gap shrinks but doesn't vanish after controlling for OpenCLIP cosine distance. That's real rigor.\n\nThe soft spot is exactly where your stress-test points. Original-over-null 2AFC is 49.8–78.1%: models separate the original from a no-op FLUX.2 regeneration. So ordinary source-vs-edit comparisons contain a detectable \"original vs generated\" signal. The paper's artifact control — null-over-semantic with both candidates generated — is only reported as a micro-average over all 2,137 edits, not broken down by family. That leaves the central \"attribute hardest\" hierarchy unverified in artifact-controlled form. Distance balancing partially addresses visual magnitude, but a family-specific regeneration artifact is still a plausible confound. The authors should either release per-family null-vs-semantic numbers or soften the hierarchy claim. Minor points: one source image per concept, one editing model, and family sizes range from 179 to 979; the balanced aggregation helps, so these are secondary.\n\nOverall, this is a serious benchmark paper. The core finding — standard retrieval does not transfer to edit-pool discrimination — holds up across eight models and survives multiple controls. Your conditional verdict is fair. Anyone working on EEG-image retrieval or brain-decoding evaluation should engage with it, and it deserves peer review. I'd send it to a competent referee and expect revision rather than rejection.","headline":"Genuinely useful controlled-edit benchmark for EEG-image retrieval, with a solid central finding, but the family-level difficulty hierarchy is not yet artifact-controlled.","tokens_in":21207,"tokens_out":2777,"would_cite":true,"duration_ms":26685,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model that scores high on standard EEG-to-image retrieval can still fail to distinguish the viewed image from controlled edits of that image, and fine-grained attribute edits are the hardest to catch.","keywords":["EEG-to-image retrieval","diagnostic benchmark","image editing","object attributes","visual information","THINGS-EEG2","two-alternative forced choice","brain decoding"],"falsifier":"Curate a new set of edits where every family is matched on measured visual distance (for example, small identity flips versus large color changes) and check whether attribute edits still underperform identity edits; if the gap collapses, the hierarchy is an artifact of edit magnitude. A second check is to record EEG from subjects who view the edited images and compare human discrimination accuracy with model 2AFC accuracy.","tokens_in":20082,"feed_emoji":"🧠","tokens_out":5503,"duration_ms":52940,"temperature":0.7,"pith_summary":"EEG-EditBench tests whether EEG-to-image retrieval models actually register the visual details of what a person viewed. The authors construct 2,137 quality-controlled edits of 200 test images, changing object identity, attributes, background, or object presence, and evaluate eight retrieval models. They find that standard 200-way retrieval accuracy does not transfer to deciding between an original image and its own edited variants: edit-pool top-1 accuracy is much lower than standard accuracy, and attribute changes are the hardest to detect. If this is right, aggregate retrieval results have been over-reading what visual information EEG-image models preserve, and edit-based tests should accompany standard metrics.","feed_headline":"Strong retrieval scores don't prove EEG models see the details","feed_subtitle":"Controlled image edits reveal that standard 200-way accuracy oversells which visual details EEG models decode.","key_machinery":"The central object is EEG-EditBench itself: 2,137 edited variants of 200 source images, organized into four edit families (object identity, attribute, background, removal), with each edit reviewed by independent human judges. Evaluation uses two controlled scores: per-family two-alternative forced choice (does the source image outscore one edited variant?) and edit-pool Top-1 (does the source image outrank all edited variants derived from it?). A null-edit control, where the editor is asked to reproduce the image without semantic change, provides a reference for image-generation artifacts.","core_discovery":"The paper's central discovery is that a model's ability to retrieve a viewed image from 200 unrelated concepts is not evidence about its ability to distinguish that image from controlled variants of the same scene. Across eight models, standard 200-way Top-1 ranges from 19.5% to 75.2%, while edit-pool Top-1 (ranking the source above all its same-source edits) falls to 14.0%–46.3%. Per-family two-alternative forced choice shows object identity and removal edits are handled well (79.1%–95.6% and 83.6%–97.7% respectively), background edits are middling, and attribute edits are hardest (55.8%–80.5%). The authors interpret the pattern as showing that current EEG-image representations preserve coa","pith_inferences":["Because the paper's own null-edit control shows that regeneration alone changes image representations enough for models to notice, part of the edit-difficulty hierarchy may be driven by editor artifacts rather than by the targeted visual factor; a direct test is to build edits matched on measured visual distance across families and see whether the attribute gap persists.","The paper notes that edited images have no corresponding EEG recordings, so the benchmark tests the complete retrieval system, not the brain's response to edited stimuli; collecting such recordings would connect the hierarchy to human visual discriminability.","If the attribute-blindness pattern generalizes, the most practical next step is not more of the same retrieval training but objective functions that force attribute-level discrimination, evaluated by whether edit-pool scores improve.","Mismatched-EEG baselines reveal that models also rely on image-side priors independent of the neural signal, so future edit-based benchmarks should routinely report such baselines before interpreting pairwise scores as evidence about brain alignment."],"forward_implications":["High 200-way retrieval accuracy can coexist with poor fine-grained edit discrimination, so published retrieval scores should not be read as evidence about which visual details are encoded.","Across all eight evaluated models, object identity and object presence are distinguished more reliably than attributes; within attributes, shape or size edits are the hardest and color edits the easiest.","Semantic proximity drives identity-edit difficulty: far replacements are easier to detect than near or medium ones.","The same attribute edit varies in difficulty across object categories, so a model's visual sensitivity is not uniform over content.","Edited variants can serve as structured hard negatives for training EEG-image models that retain more precise and interpretable visual information."],"fun_headline_variants":["EEG retrieval scores oversell what models truly see","Fine-grained attributes trip up EEG-image retrieval models","Same-scene edits expose blind spots in EEG retrieval","Standard accuracy hides EEG models' weak attribute grasp","Attribute edits are the toughest test for EEG-image models"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that each retained edit changes only the intended visual factor (identity, attribute, background, or removal) and preserves everything else; if regeneration artifacts co-vary with edit family, the observed difficulty hierarchy could come from the editor rather than from what the EEG-image model preserves.","fun_headline_variants_meta":{"raw":{"variants":["EEG retrieval scores oversell what models truly see","Fine-grained attributes trip up EEG-image retrieval models","Same-scene edits expose blind spots in EEG retrieval","Standard accuracy hides EEG models' weak attribute grasp","Attribute edits are the toughest test for EEG-image models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000543,"raw_usage":{"total_tokens":2433,"prompt_tokens":738,"completion_tokens":1695,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":1635}},"tokens_in":482,"tokens_out":1695,"duration_ms":11654,"temperature":1.0,"reasoning_tokens":1635,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:10:28.262320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Curate a new set of edits where every family is matched on measured visual distance (for example, small identity flips versus large color changes) and check whether attribute edits still underperform identity edits; if the gap collapses, the hierarchy is an artifact of edit magnitude. A second check is to record EEG from subjects who view the edited images and compare human discrimination accuracy with model 2AFC accuracy.","supporting_citations":[],"review_version":1}