{"id":"35591335-6fb0-494e-8fc1-b94c76858b80","arxiv_id":"2608.05702","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SEAM encodes regional model explanations as a sheaf and uses the coboundary operator to turn overlap disagreements into a channel-resolved obstruction that can localize and test repair hypotheses.","lead":"SEAM is a framework that checks whether local explanations from scientific models can be glued into one globally consistent account, comparing neighboring regions on their overlaps and reporting any mismatch as a channel-resolved obstruction. It matters because locally accurate models can still disagree across region boundaries, and SEAM offers a structured way to detect, localize, and test repair hypotheses for such disagreements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exact-feasibility verdicts are brittle: SEAM-Ω's formal refute/retain branch requires exact zero residuals, which real noisy explanations will almost never satisfy, so the decisive verdict defaults to 'unresolved' outside noiseless synthetic experiments.","rationale":"The reader's conditional verdict is appropriate. My stress-test identified a different load-bearing soft spot than the reader's pairwise/A0 concern: the formal verdict depends on exact zero residuals, which are structurally fragile for real, noisy, discretized explanations. The paper is internally consistent — Definition 2.5 explicitly separates exact, tolerance, and soft branches — and it honestly labels soft records as empirical attribution. Precisely because of that honesty, the advertised scientific verdict of Definition 2.5 is almost never attainable in practice: exact feasibility requires the obstruction to lie in the image of the restricted coboundary, and under A0 for single-channel budgets this requires all non-target channel components to vanish. Real data will not satisfy this. The paper's most impressive recovery and attribution experiments operate in the soft or projected branch rather than the exact branch, so the exact theorems are not the operative machinery in the noisy regime. This does not falsify the central claim; it scopes it. The concrete perturbation test would show whether even a 1e-8 perturbation flips a unique retained account to unresolved; if it does, the exact branch cannot be the basis for real-world attribution, strengthening the need for the conditional verdict and for reporting tolerance-based verdicts as primary rather than engineering fallbacks. No change to the reader's CONDITIONAL verdict is required; the concern reinforces it.","tokens_in":43901,"tokens_out":16198,"duration_ms":188360,"concrete_test":"Run Algorithm 2 with the three standard single-channel budgets on the UCI household-power cross-framework audit (Section 9.2.3) and record the exact-feasible proper budget set; under noisy real data this set will almost certainly be empty, making the formal verdict 'unresolved'. Then perform a controlled perturbation test on a two-region cover with state and closure channels: construct an obstruction concentrated in the closure channel with the state component exactly zero, so the closure-revision budget is exact-feasible; add independent Gaussian noise of scale 1e-8 to the state vectors and recompute. If the closure-revision budget flips to infeasible, exact feasibility is not robust to infinitesimal noise, settling whether the refute/retain branch can support real-world attributions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central budgeted-intervention claim (Abstract; Section 4.2; Definition 2.5; Algorithm 2) is that exact feasibility omega in im(DP) refutes or retains a declared account. Under Assumption A0, a single-channel budget is feasible only if the obstruction components in all other channels vanish exactly. Real scientific explanations are noisy and discretized, so every channel will generically have nonzero components; the proper feasible set will be empty, and the formal verdict becomes 'unresolved by the declared specific accounts'. The paper's own hidden-source recovery (Section 9.3.1) illustrates this: it reports a state-channel norm of 0.005873 and a closure-channel norm of 2.020986, so the closure-revision budget is infeasible under exact semantics; the paper switches to a projected reconstruction with residual below 5%. Similarly, the data-physics conflict attribution (Section 9.4) uses soft squared intervention norms because exact single-channel hard costs are +infinity whenever non-target components are nonzero. Thus the advertised hypothesis-testing verdict is operational only when the obstruction is exactly channel-concentrated, a condition that noiseless synthetic experiments are engineered to satisfy but real data essentially never will. Theorems 1 and 2 are correct conditional statements, but the empirical support does not demonstrate that the decisive exact branch survives contact with real data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Scientific Explanation-Admissibility Machines (SEAM), a framework for auditing whether locally plausible scientific explanations can be assembled into a globally consistent account. The concrete instantiation SEAM-Omega represents each region by a finite-dimensional explanation stalk with state, closure, observation, and optional metadata channels; restriction maps compare neighboring explanations on pairwise overlaps; the coboundary matrix D assembles these restrictions; and the obstruction omega = D s measures the failure of global admissibility. The paper develops a channel-resolved diagnosis of omega, a budgeted intervention framework in which declared scientific hypotheses are represented as projection matrices P and tested by exact feasibility of omega in im(DP), an identifiability analysis based on ker D intersect ker O, and a streaming monitoring extension. The theoretical results include a closed-form minimum-cost intervention theorem (Theorem 1), identifiability dimension formulas (Proposition 7.2), closure-restricted recoverability (Corollary 7.3), and a conservation-contract detectability bound (Theorem 2). The empirical section reports nineteen experiments spanning synthetic PDEs, Fourier neural operator monitoring, four public datasets, and synthetic financial and industrial systems.","tokens_in":44198,"tokens_out":4139,"duration_ms":48966,"significance":"If the central claims hold, the paper offers a genuinely useful formal object: a computable, channel-resolved obstruction that separates local accuracy from global explainability, together with an auditable procedure for testing competing repair hypotheses. The theoretical core is straightforward and mostly correct: the definitions of the coboundary and obstruction are standard, the pseudoinverse-based feasibility and minimum-cost formulas in Theorem 1 are properly derived, and the identifiability and closure-recovery statements follow from elementary linear algebra. The paper is also commendably explicit about its assumptions: Assumption A0 (block-diagonal channel restrictions) and Assumption A3 (the conditions of Theorem 2) are stated, and Section 11.1 openly declares the pairwise-overlap limitation. The experimental protocol is unusually detailed about seeds, tolerances, and data provenance, and the zero-floor control and monotonicity sweep are useful sanity checks.","major_comments":[{"comment":"The hidden-source recovery experiment is circular as a validation of closure recoverability. The setup states that 'the local explanation emits a closure-summary block measuring the missing-source residual on the overlap' and that 'this block is audited through the closure restriction.' In other words, the injected source is encoded in the very channel whose recovery is then reported. It is therefore not surprising that the closure channel contributes 99.9992% of ||omega||^2 and that the closure repair succeeds. The experiment would be informative only if the backend's closure block were not defined as the missing-source residual, for example if the source were injected into the state dynamics while the closure block is learned or is a generic closure parameter. As written, the claim 'missing closure is recoverable' is built into the data-generation step rather than demonstrated by the audit.","section":"§9.3.1"},{"comment":"The advertised decisive branch of the framework—exact feasibility, omega in im(DP), refuting or retaining a declared account—is fragile under realistic noise and discretization. Under Assumption A0, a single-channel hard budget is feasible only when the obstruction components in all other channels vanish exactly. Real scientific explanations are noisy, discretized, and approximate, so every channel generically carries a nonzero component; the formal verdict then falls into the 'unresolved by the declared specific accounts' branch. The paper's own experiments illustrate this. In §9.3.1, the saved obstruction has a state-channel norm of 0.005873 alongside the closure-channel norm of 2.020986, so the exact closure-revision budget is infeasible by the strict definition; the paper switches to a projected reconstruction with residual below 5%. In §9.4, the data–physics conflict attribution uses soft squared intervention norms because exact single-channel hard costs are +infinity whenever non-target components are nonzero. The manuscript does not report a single real-data or realistically noisy experiment in which the exact branch returns a definitive retain/refute verdict. This does not invalidate the theorems, which are correct conditional statements, but it means the central operational claim of the abstract and Section 1 is not supported by the evidence. The paper should either demonstrate the exact branch on data with realistic noise, or explicitly and prominently characterize the exact branch as a noiseless/synthetic-only guarantee and present the soft records as the primary practical output.","section":"Definition 2.5; Algorithm 2; §9.3.1; §9.4"},{"comment":"The FNO OOD monitoring experiment shows a strong correlation between ||omega(t)|| and the reference L2 error, but this is an uncalibrated empirical association, not a detection guarantee, and the exact feasibility branch plays no role in the monitoring use case. The paper does state in §6 that 'SEAM does not assign a universal alarm threshold' and in §9.8.2 that the experiment 'does not establish a calibrated probabilistic detector.' That honesty is appreciated. Still, the abstract's claim that 'SEAM detects incompatible explanations even when local predictions are accurate' is repeatedly supported by correlation-style evidence rather than by a decision procedure with controlled error rates. A revision should either add threshold-based detection evaluation (e.g., ROC or precision-recall against injected shifts) or consistently phrase the monitoring claim as 'correlates with' rather than 'detects.'","section":"§9.8.2 and §6"},{"comment":"The framework's validity depends on Assumption A0 and on the modeling premise that a backend's local explanation can be faithfully represented as a finite-dimensional vector with block-diagonal linear restriction maps. The paper declares this limitation in Section 11.1, which is appropriate. However, the consequences for the channel diagnostics are stronger than the discussion suggests: if a backend's explanation contains cross-channel coupling, or if inconsistencies arise only in triple-overlap interactions, then every channel-dominance report and every budgeted verdict built on the 1-skeleton and block-diagonal restrictions can be incomplete or misleading. The manuscript should state explicitly, in the introduction and in the interpretation guide, that all channel diagnoses and budget verdicts are conditional on these representational choices, not merely on the pairwise-overlap approximation.","section":"Section 3.4 and Section 11.1"}],"minor_comments":[{"comment":"The definition of the hard residual writes r_hard_P(omega) = omega - (DP)(DP)^+ omega in ker(DP)^\top; the notation 'ker(DP)^\top' is nonstandard. It should be the orthogonal complement of the range of DP, or equivalently ker((DP)^\top). This is a notation issue, not a mathematical error.","section":"Definition 3.9"},{"comment":"The piecewise verdict in Definition 2.5 uses tau_zero for the global admissibility branch, but no guidance is given for choosing tau_zero in practice, even though the later experiments fix it at 10^-6. Since the exact branch is so sensitive to tiny nonzero components, a short paragraph on how to set tau_zero relative to discretization error or sensor noise would materially improve the operational usefulness.","section":"Section 2.3, Definition 2.5"},{"comment":"The seed protocol is described as 'a fixed set of n = 5 seeds,' and deterministic experiments are said to be 'executed across the same schedule.' This is acceptable, but the notation '0.0266±0.0000' for a deterministic result is confusing; it would be clearer to report deterministic results as exact values and reserve plus-minus notation for genuinely stochastic quantities.","section":"Appendix E.1"},{"comment":"The random-sinusoid Burgers stress test reports ||omega|| = 0.4754±0.0000, but the reader is not told what the obstruction norm would be if the same generator produced a globally admissible family. Without an admissibility control for the same random-sinusoid family, the experiment demonstrates that omega is nonzero but not that it is informative as a stress test.","section":"Section 9.7.1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical backbone is sound and the paper is refreshingly explicit about its assumptions and limitations. My main concern is the gap between the advertised exact-feasibility verdict and the actual experimental demonstrations, which almost always fall back to projected or soft records. The hidden-source recovery experiment is also circular as currently constructed. I would support publication after the authors either demonstrate the exact branch on realistic noisy data or explicitly reframe the exact branch as a noiseless guarantee and promote the residual-aware records to the central practical output. The pairwise-overlap and block-diagonal limitations are declared, so I do not treat them as fatal, but the manuscript should connect those limitations to the validity of the channel diagnostics more forcefully."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SEAM is a serious, well-built framework for auditing whether locally accurate scientific explanations glue into a globally consistent account. The core construction is clean: incident explanations are compared on pairwise overlaps, disagreement is encoded as omega = D s, and channel decomposition plus budgeted pseudo-inverse repairs turn each causal story into a feasibility test. The sheaf framing is a genuine new application in scientific ML, and the paper is unusually careful about what its verdicts can and cannot say.\n\nThe theorems are correct as stated, the definitions are coherent, and the limitations are mostly declared up front: pairwise overlaps only (Section 11.1), catalog dependence, and no claim that obstruction norm measures absolute accuracy. The zero-floor negative controls and the FNO OOD correlation (r=0.995 on Burgers) are nice evidence that the machinery behaves.\n\nThe main soft spot is exactly what the stress-test note flags: the formal verdict in Definition 2.5 is driven by exact feasibility, and real noisy explanations will almost never lie exactly in a budget's permitted image. So the decisive refute/retain branch is operational only in synthetic channel-concentrated cases; on real data the framework silently falls back to soft records with residuals. The paper labels these records honestly, but the headline claim that SEAM 'tests competing accounts' is stronger than what a practitioner gets on noisy data. The paper's own hidden-source recovery uses a projected reconstruction with residual below 5%, and the conflict-attribution study uses soft norms, so the exact branch is never demonstrated on a realistic noisy example. That's a genuine gap, not a manufactured one.\n\nThe empirical validation is also somewhat self-confirming: the closure-summary block is defined as the missing-source residual, so recovering it is in part an algebraic check, and most experiments are engineered so one channel dominates. No code or data is released, which matters for an audit-style framework.\n\nBottom line: this deserves a serious referee. It is a thoughtful, mathematically grounded contribution that a good editor should send to review, with the expectation of revision. The authors should be asked to release code and data and to show a decisive verdict on a noisy, non-synthetic case where the exact branch, or an explicitly defined tolerance branch, actually discriminates between accounts.","headline":"A coherent sheaf-based audit for global consistency of scientific explanations, with an exact-feasibility verdict that is honest but narrowly applicable to noiseless channel-concentrated cases.","tokens_in":44642,"tokens_out":2838,"would_cite":false,"duration_ms":30632,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["55N30","15A09"],"pacs":[],"model":"deepseek-v4-flash","headline":"Explanation-admissibility is exactly computable as a single defect vector, and SEAM-Ω turns that vector into a scientific verdict.","keywords":["explanation-admissibility","cellular sheaves","scientific machine learning","local-to-global consistency","budgeted intervention","identifiability","Fourier neural operator","obstruction-guided diagnosis"],"falsifier":"Run the same backend through a three-region cover whose explanations agree on every pair but violate a constraint that only appears when all three regions are considered together, and with all restrictions block-diagonal; if SEAM-Ω reports $\\omega = 0$ while the triple-overlap mismatch exists, the pairwise 1-skeleton verdict is incomplete.","tokens_in":43693,"feed_emoji":"🧩","tokens_out":6121,"duration_ms":58457,"temperature":0.7,"pith_summary":"The paper sets out to make a new scientific object computable: explanation-admissibility, the property that a family of locally generated explanations can be assembled into one globally coherent account. It claims that this property is independent of local accuracy—individual regions can pass every local test while the family as a whole cannot be glued together. SEAM-Ω is the finite linear instantiation: each region's explanation is a vector with state, closure, and observation channels, neighboring explanations are compared on their overlaps through linear restriction maps, and the disagreement is condensed into a single obstruction $\\omega = D s$. Nonzero $\\omega$ certifies global inadmissibility, its channel decomposition locates the mismatch, and exact feasibility of budgeted repairs refutes or retains competing causal accounts. If the claim holds, scientific machine learning gains an audit that detects inconsistent explanations even when every local model scores well.","feed_headline":"One defect vector exposes globally incoherent explanations","feed_subtitle":"Locally accurate models can disagree on shared boundaries; SEAM turns that disagreement into a channel-by-channel diagnosis.","key_machinery":"The central object is the finite explanation sheaf and its 1-cochain obstruction $\\omega = D s$. Each region carries a stalk partitioned into state, closure, and observation channels, with optional contract metadata; each overlap carries an overlap stalk, and the hand-crafted linear restriction maps $\\rho_{i,ij}$ extract the quantities that must agree. Under Assumption A0 the restrictions are block-diagonal per channel, so the channel decomposition of $\\omega$ commutes with the coboundary. The argument is carried by exact conditions: admissibility is $\\omega = 0$, a budgeted hypothesis $P$ survives exactly when $\\omega \\in \\operatorname{im}(D P)$, and the closed-form minimum-cost repair uses the restricted pseudoinverse $A_P^+$ of Theorem 1, with the blind admissible subspace $\\ker D \\cap \\ker O$ recording what observations cannot resolve.","core_discovery":"The central claim, stated on the paper's own terms, is that the local-to-global consistency of scientific explanations is exactly certifiable as linear algebra. Given a cover of regions, stacked explanations $s$, and restriction maps assembled into a coboundary matrix $D$, the obstruction $\\omega = D s$ vanishes if and only if neighboring restrictions agree on every overlap, so $\\omega$ is the exact certificate of admissibility rather than a heuristic score. When $\\omega \\neq 0$, the block-diagonal channel structure of the restrictions splits $\\omega$ into state, closure, and observation components; the budgeted intervention theorem then tests each declared account by checking whether the revisions that account permits can remove the whole obstruction, i.e., whether $\\omega \\in \\operatorname{im}(D P)$, with the minimum-cost repair given in closed form by a pseudoinverse. The paper also claims that identifiability separates inconsistency from directions the observations cannot see, and that the streaming obstruction tracks learned generators under distribution shift, as evidenced by nineteen experiments including the Fourier neural operator monitoring study.","pith_inferences":["Editorial inference: the pairwise-overlap 1-skeleton is a declared limitation, so a natural next test is whether triple-overlap obstructions capture failures that pairwise agreement misses; the paper's own Section 11.1 leaves this open.","Editorial inference: because the restriction maps are linear and the sheaf is finite, the construction generalizes immediately to heterogeneous backends (solvers, regressors, neural operators) once a common channel schema is fixed; a reader could apply it to multimodal sensor fusion.","Editorial inference: the minimum-cost repair cost, while not a cross-account ranking, could serve as a quantitative measure of how much revision a surviving hypothesis demands, enabling comparisons of required interventions across repeated audits.","Editorial inference: the identifiability subspace suggests a data-acquisition guideline: add observations that shrink $\\ker D \\cap \\ker O$ before trust is placed in closure attributions."],"forward_implications":["A SEAM audit can certify global admissibility of a deployed scientific model family in closed form, at the cost of one matrix-vector product once $D$ is assembled.","Local validation, benchmark splits, and residual checks are no longer sufficient evidence of coherence; systems must also pass the overlap restriction test.","Reported channel dominance converts a vague 'something is inconsistent' into a concrete localization: state, closure, or observation channel, and which overlaps carry the defect.","Budgeted intervention turns diagnostic disagreement into a hypothesis test: a declared account is refuted if its permitted revisions cannot remove the whole obstruction, and priced by its minimum feasible repair cost.","The streaming obstruction $\\omega(t)$ provides a distribution-shift monitor for learned operators that correlates with prediction error and is cheaper than ensemble-variance baselines."],"supporting_citations":[{"why":"Supplies the sheaf and cosheaf foundation from which the finite explanation sheaf and its nerve complex are drawn.","marker":"Curry (2014)"},{"why":"Introduces the global-section viewpoint that underlies the definition of explanation-admissibility.","marker":"Ghrist (2014)"},{"why":"Systematizes cellular sheaves, providing the stalk-and-restriction machinery that SEAM-Ω instantiates.","marker":"Robinson (2014)"},{"why":"Defines the Moore–Penrose pseudoinverse used in the closed-form minimum-cost intervention formula and closure recovery.","marker":"Penrose (1955)"},{"why":"Provides the Fourier neural operator architecture used in the out-of-distribution monitoring experiments.","marker":"Li et al. (2021)"}],"fun_headline_variants":["SEAM's linear algebra pinpoints global explanation failures","Local accuracy isn't enough: SEAM certifies global coherence","From local checks to global admissibility: SEAM computes it","Channel-resolved obstructions: SEAM's exact inconsistency test","Catch hidden explanation conflicts with SEAM's exact audit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework stands or falls on the modeling premise that a real backend's local explanation can be faithfully represented as a finite vector with state, closure, and observation blocks and that all relevant inconsistencies appear on pairwise overlaps, a restriction the paper itself declares.","fun_headline_variants_meta":{"raw":{"variants":["SEAM's linear algebra pinpoints global explanation failures","Local accuracy isn't enough: SEAM certifies global coherence","From local checks to global admissibility: SEAM computes it","Channel-resolved obstructions: SEAM's exact inconsistency test","Catch hidden explanation conflicts with SEAM's exact audit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000746,"raw_usage":{"total_tokens":3367,"prompt_tokens":1028,"completion_tokens":2339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":2256}},"tokens_in":644,"tokens_out":2339,"duration_ms":18103,"temperature":1.0,"reasoning_tokens":2256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:44:54.837552+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same backend through a three-region cover whose explanations agree on every pair but violate a constraint that only appears when all three regions are considered together, and with all restrictions block-diagonal; if SEAM-Ω reports $\\omega = 0$ while the triple-overlap mismatch exists, the pairwise 1-skeleton verdict is incomplete.","supporting_citations":[{"cited_title":", year 2014","cited_arxiv_id":null,"evidence_quote":"Supplies the sheaf and cosheaf foundation from which the finite explanation sheaf and its nerve complex are drawn."},{"cited_title":", year 2014","cited_arxiv_id":null,"evidence_quote":"Introduces the global-section viewpoint that underlies the definition of explanation-admissibility."},{"cited_title":", year 2014","cited_arxiv_id":null,"evidence_quote":"Systematizes cellular sheaves, providing the stalk-and-restriction machinery that SEAM-Ω instantiates."},{"cited_title":", year 1955","cited_arxiv_id":null,"evidence_quote":"Defines the Moore–Penrose pseudoinverse used in the closed-form minimum-cost intervention formula and closure recovery."}],"review_version":1}