{"id":"14ae2c38-2066-43f5-acdf-682db835d7f3","arxiv_id":"2607.25387","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A learned restriction map can be heavily used by a trained sheaf GNN checkpoint yet be replaceable after retraining; only Roman-Empire among five benchmarks shows persistent edge-assignment value.","lead":"This paper separates \"a sheaf GNN checkpoint uses its learned edge maps\" from \"the task actually needs those maps,\" by intervening on a fixed model versus retraining matched alternatives. On four of five benchmark datasets, retraining with shuffled or shared maps recovers full performance; only Roman-Empire keeps a real edge-assignment advantage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Roman-Empire's 'unreplaced' verdict is indexed to a finite control suite; an untested edge-gating control could close the gap, so 'indispensable' overreaches.","rationale":"I could not identify an internal mathematical error: Theorem 1 and the frame boundary are correct for the stated linear model, and the empirical audit is carefully paired. The main risk is interpretive: 'indispensable' is protocol-relative, exactly as the reader's weakest_assumption states. The proposed Scalar-Gate-F control is a direct test of whether an unlisted replacement changes the Roman-Empire verdict. It preserves operator, splits, budget, and validation rule while changing only the mechanism that uses edge-indexed parameters, so it isolates whether the unreplaced advantage is specific to matrix transport or merely to persistent edge indexing. Because the central existence claim is already supported by four datasets even if Roman-Empire later falls, the reader's CONDITIONAL verdict remains appropriate; no change is needed.","tokens_in":17100,"tokens_out":18533,"duration_ms":192650,"concrete_test":"Implement a new control family, Scalar-Gate-F, under the same protocol: replace the d×d incidence-map path with one learned scalar coefficient per directed edge, computed from endpoint embeddings by a small MLP, plus node-wise adapters to match the Full parameter budget. Train on the ten official splits with the same LR grid, early stopping, and paired evaluation. If Scalar-Gate-F still trails Full on Roman-Empire with a CI excluding zero, the 'unreplaced' classification is robust to a broader class of edge-dependent replacements; if it closes the gap, the conclusion should be weakened to 'unreplaced only within the tested map-removal controls.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has an existence component—learned maps can be checkpoint-reliant yet replaceable—and the four non-Roman datasets support it well: Resampled-F matches Full with the same parameter count and per-forward map multiset. The load-bearing weakness is the stronger reading of the conclusion, 'without constituting indispensable edge geometry,' and the Roman-Empire 'unreplaced' categorisation, because those depend on the pre-specified control suite C in Definition 2/Eq. (4) being exhaustive in a way that matters. Every control that removes native assignment either destroys edge variation (Shared-F/C), fixes a wrong stable assignment (Fixed-F), or makes assignment stochastic (Resampled-F). None tests an edge-dependent but non-transport mechanism (e.g., scalar edge gates or attention) that could absorb the task while retaining persistent edge indexing. If such a control closed Roman-Empire's +.0675 margin, the 'suite-unreplaced' verdict would reverse; if it also recovered the other four, the claim would still hold but only for a narrower sense of 'transport.' The paper acknowledges this scope in §6 ('indexed by the operator, training budget, validation rule, and control suite'), but the abstract's 'indispensable edge geometry' is stronger than what a finite suite can establish.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper distinguishes three evidential claims about learned restriction maps in sheaf GNNs: map movement (training history), checkpoint reliance (post-hoc intervention on a fixed checkpoint), and protocol-relative replacement (matched retraining against control families). It contributes a task-null subspace theorem (Theorem 1) showing that labels constrain only the transported classifier directions, leaving d^2−d invisible degrees of freedom per full d×d map; an exact frame model (Theorem 2, Corollary 1) giving a reliance–replacement boundary; a label-only experiment realizing the predicted separation; and a paired audit of public NSD/DNSD/DSNN implementations. The audit finds that all five DNSD full checkpoints rely on their learned maps, but after retraining, assignment-breaking or shared-map controls recover Full performance on four benchmarks, while Roman-Empire retains a 0.0675 advantage over continually resampled assignment and a 0.0391 advantage over a parameter-matched shared map across ten official splits. The paper concludes that a learned map can govern a fitted computation without constituting indispensable edge geometry, and that claims of learned transport should pair checkpoint interventions with matched retraining.","tokens_in":17451,"tokens_out":14389,"duration_ms":181009,"significance":"If the results hold, the paper makes a valuable and timely methodological point: post-hoc map ablations or parameter-movement diagnostics do not establish task-level necessity for learned sheaf transport. The theoretical results are clean within their stated linear/Gaussian models, are parameter-free, and are verified by a correspondence-supervised calibration and a label-only experiment. The real-graph audit is carefully paired: it uses official splits, confidence intervals, learning-rate robustness checks, extra optimization seeds, an exact sign-flip test, and intervention verification. The conceptual ladder in Figure 1 and the explicit protocol-relative definition of replacement are useful for future work. The main limitation—that the 'unreplaced' verdict is relative to the pre-specified control suite—is acknowledged in §6, so it is a scope condition rather than an internal inconsistency.","major_comments":[],"minor_comments":[{"comment":"The phrases 'unreplaced transport regimes' and 'indispensable edge geometry' are stronger than what Definition 2/Eq. (4) and the finite control suite can establish. The paper's own §6 says 'suite-unreplaced'; the abstract should carry the same qualifier, e.g., 'not replaced by the tested control suite' or 'not indispensable under the tested operators, budgets, and controls.'","section":"Abstract and §7"},{"comment":"The quantities R_I and N_I are used in Eq. (13) but are not explicitly defined in the main text. Please define them in terms of the three accuracies in Theorem 2 (e.g., R_I = A_oracle − A_fixed^I and N_I = A_oracle − A_retrain^I) before the corollary is stated.","section":"Corollary 1"},{"comment":"The label-only experiment is a qualitative realization of the separation, but a reader may expect the numbers to match Theorem 2's A_fixed^I. It would help to state explicitly that Theorem 2 pertains to the oracle-aligned classifier, whereas the jointly trained label-only model can select a non-optimal relied-upon factorization (as in Theorem 1), so the two are not the same quantities.","section":"Section 4"},{"comment":"The text says 'All 210 sign assignments are enumerated.' For ten paired splits, the number of sign assignments is 2^10 = 1,024. Please correct the number or clarify if a subset was enumerated.","section":"Appendix E.3"},{"comment":"The statement says code and run-level records 'will be released publicly in a forthcoming project release' rather than being currently available. Please provide an anonymized repository or a clear availability statement at acceptance; the audit's reproducibility depends on the exact control implementations.","section":"Reproducibility Statement"}],"recommendation":"minor_revision","confidential_remarks":"The paper is technically sound within its stated scope. The central distinction between checkpoint reliance and protocol-relative replacement is well supported by theory and experiments. The only substantive concern is that the abstract and conclusion use absolute-sounding language ('unreplaced', 'indispensable') that goes beyond the finite control suite; this is acknowledged in §6 and can be fixed with wording changes. The '210' sign-assignment typo and missing definitions in Corollary 1 are also easy local fixes. I do not see a need for additional experiments before acceptance, though an edge-gating control would strengthen the Roman-Empire claim if the authors can add it cheaply."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this paper gets the central distinction right. A post-hoc ablation on a fixed checkpoint tells you how that checkpoint is wired, not whether the task needs the removed component. The authors formalize that with two estimands and then do the retraining work.\n\nWhat's new: the d^2-d task-null theorem is simple but useful — labels only pin down transported classifier directions, leaving most matrix degrees of freedom free. The exact frame model gives a clean boundary where reliance exceeds replacement. The audit protocol — fixed vs resampled assignment, layer-shared maps, parameter-matched adapters — is the right way to test whether persistent edge geometry matters. The real-graph results are carefully paired: ten splits, confidence intervals, learning-rate robustness, extra seeds, and intervention checks. The finding that all five DNSD benchmarks rely on maps but only Roman-Empire survives retrained replacement is credible.\n\nThree things to fix. First, the abstract says 'without constituting indispensable edge geometry,' which is stronger than the evidence: replaceability is defined relative to a pre-specified control suite and budget. The paper says this in §6, but the abstract doesn't. Second, the text claims 'All 210 sign assignments are enumerated,' but with ten paired differences the exact sign-flip test has 1024 assignments. That looks like a typo, but it should be corrected because the p-values depend on it. Third, code is promised but not yet released; the reproducibility statement is a promise. Also, the d^2-d theorem is proven for the linear frame model; generalizing it in the abstract to all learned maps is fine if the wording makes the model scope explicit.\n\nNone of these break the central argument. The stress-test worry about an untested edge-gating control is real, but the paper already scopes its claim; Roman-Empire is 'suite-unreplaced,' not 'indispensable' in an absolute sense.\n\nWho should read: anyone doing interpretability or architecture attribution for GNNs, and especially people who cite sheaf GNNs as evidence their geometry is doing the work. This deserves a serious referee; the protocol is a contribution that should be reviewed and refined.\n\nSend it to review. It's a solid methodological paper with clean math and careful empirics; the fixes are minor.","headline":"Careful separation of checkpoint reliance from retrained task value; the audit is solid even if the abstract's 'indispensable' overreaches the protocol.","tokens_in":17903,"tokens_out":2624,"would_cite":true,"duration_ms":25709,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A learned restriction map can govern a fitted checkpoint while remaining replaceable after retraining, so post-hoc ablations only prove reliance, not task-level necessity.","keywords":["sheaf graph neural networks","restriction maps","checkpoint reliance","task-level value","matched retraining","protocol-relative replacement","task-null subspace","heterophilous graphs"],"falsifier":"Train a parameter-matched node-wise adapter or a more expressive message-passing head on Roman-Empire that does not use per-edge learned maps but has greater capacity than the tested controls; if it closes or reverses the .0675 advantage over Resampled-F, the claim of suite-unreplaced value on that dataset fails. Conversely, if a similar richer control closes the gaps on any of the other four datasets, the claim that they do not need persistent assignment would need revision.","tokens_in":16945,"feed_emoji":"🔄","tokens_out":3519,"duration_ms":31624,"temperature":0.7,"pith_summary":"The paper argues that evidence commonly used to claim sheaf graph neural networks discover useful edge geometry—maps moving away from identity and post-hoc replacement hurting predictions—only shows that one fitted checkpoint relies on its maps. It does not show the task needs them. The paper separates two estimands: checkpoint reliance, which intervenes on the maps of a fixed predictor, and protocol-relative replacement, which retrains matched model families that remove map capacity, edge variation, or persistent edge assignment. A task-null theorem explains why these can diverge: labels constrain only d of the d^2 directions in each full d×d map. Empirically, on four of five benchmarks, retrained controls that break or share edge assignment recover full performance; only Roman-Empire retains a persistent-assignment advantage.","feed_headline":"Retrained controls reveal most learned sheaf maps are replaceable","feed_subtitle":"Post-hoc ablation measures reliance, not necessity; matched retraining separates the two on real graphs.","key_machinery":"The central objects are the induced cross-edge transport action A_{u→v} = F_{v⊴e}^T F_{u⊴e} and the protocol-relative replacement margin N_B, with a pre-specified control suite C. The task-null theorem (B_g^T w = 0) establishes that each full d×d map has d^2 − d invisible degrees of freedom under label supervision. The exact frame model, based on the finite geometric sum z_G(η), gives a closed-form boundary where checkpoint reliance and unreplaced task value separate. These mechanisms together define the evidential ladder from map movement to post-hoc reliance to suite-unreplaced task value.","core_discovery":"The central claim is that a learned transport map can be load-bearing for a single trained checkpoint while being replaceable once the rest of the model adapts. The paper proves a task-null subspace: for a nonzero linear classifier w, any d×d map decomposes into a rank-one part w q^T / ||w||^2 that affects the prediction and a d^2 − d dimensional subspace orthogonal to w that labels never identify. It constructs an exact frame model where identity intervention hurts a fixed classifier while a retrained identity classifier recovers accuracy, giving the boundary between reliance and unreplaced value. On real graphs, all full checkpoints lose score when their maps are intervened on, but after m","pith_inferences":["The reliance-replacement distinction likely extends beyond sheaf GNNs to any learned structural component (attention patterns, routing modules) where post-hoc ablation is used as evidence of necessity; this paper provides a template for such attributions.","The Roman-Empire advantage is defined against the tested control suite and training budget; a richer head or a different propagation operator could in principle absorb the transport, so that dataset-level verdict remains provisional.","The frame model's exact boundary could be adapted to nonlinear classifiers or multi-layer settings to test whether the separation persists beyond the linear single-layer case.","Reporting the first control that closes the gap, as this paper does, could become a standard attribution protocol for graph neural network components."],"forward_implications":["Post-hoc ablations of learned restriction maps should be reported as reliance tests, not as evidence of learned edge geometry.","Claims that a sheaf operator 'discovers' useful transport need matched retraining against controls that can absorb the transport, such as identity, shared maps, or resampled assignment.","On Amazon-Ratings, Questions, Tolokers, and Minesweeper, persistent edge assignment is not needed for task-level performance under the tested protocols.","On Roman-Empire, structured edge-varying actions retain value even after full retraining with randomized or shared-map controls.","The d^2 − d task-null subspace implies label supervision alone cannot identify full restriction maps; map recovery requires additional supervision or inductive bias."],"fun_headline_variants":["Most learned sheaf maps are replaceable after retraining","Retraining shows learned transport is not indispensable edge geometry","Checkpoint reliance isn't task value in sheaf GNNs","Learned sheaf maps: load-bearing but replaceable","Fixed-checkpoint reliance doesn't imply indispensable value"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The verdict that four datasets do not need persistent edge assignment rests on the assumption that the chosen control suite (identity, diagonal, capacity-matched adapters, fixed and resampled shuffles, layer-shared maps) exhausts the ways a retrained model could absorb the transport; an unlisted control, more training, or a richer head could change the margins.","fun_headline_variants_meta":{"raw":{"variants":["Most learned sheaf maps are replaceable after retraining","Retraining shows learned transport is not indispensable edge geometry","Checkpoint reliance isn't task value in sheaf GNNs","Learned sheaf maps: load-bearing but replaceable","Fixed-checkpoint reliance doesn't imply indispensable value"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3018,"prompt_tokens":810,"completion_tokens":2208,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2142}},"tokens_in":554,"tokens_out":2208,"duration_ms":27682,"temperature":1.0,"reasoning_tokens":2142,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:25:58.808037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a parameter-matched node-wise adapter or a more expressive message-passing head on Roman-Empire that does not use per-edge learned maps but has greater capacity than the tested controls; if it closes or reverses the .0675 advantage over Resampled-F, the claim of suite-unreplaced value on that dataset fails. Conversely, if a similar richer control closes the gaps on any of the other four datasets, the claim that they do not need persistent assignment would need revision.","supporting_citations":[],"review_version":2}