{"id":"c357e8b5-6372-4229-98bf-ba654ebe6621","arxiv_id":"2608.13228","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A capability-sheaf model with an exact CSP control proves that quotienting hidden-state nuisance improves a controlled harness-repair task, while a real-repository test finds no cohomological advantage over matched noncohomological selectors.","lead":"This paper models failures in AI agent 'harnesses' using sheaf theory, a branch of math about gluing local information into a consistent whole. It shows the method fixes a synthetic hidden-state problem but fails to beat simpler baselines on real software-repair benchmarks, a rare and useful negative result.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the real-repo evidence-model dependence is acknowledged and the central claim is explicitly bounded to the tested complex.","rationale":"The paper is unusually disciplined: it reports negative results, seals confirmatory splits, and draws no broad superiority claim. The reader's weakest assumption (evidence-model dependence) is real, but Section 8 already flags it and the conclusion is worded as 'tested obligation-and-risk complex is too weak,' so the assumption cannot overturn the stated claim. The controlled experiment is internally consistent, the algebraic identifiability correction is self-contained, and the negative real-repo result is not overgeneralized. Verdict remains accept.","tokens_in":15226,"tokens_out":10798,"duration_ms":116955,"concrete_test":"Re-run the frozen hidden-state analysis from the released artifact (data/latent_quotient/raw) and recompute the paired sign test over the 20 task clusters; confirm every raw-minus-quotient difference is exactly +1 and p=2^-20, with the aligned-interior ablation giving zero difference in all clusters.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I reviewed the central claims in good faith. The controlled hidden-state result is supported by a preregistered 20-cluster design, a negative control (aligned interior removes the gap), and an exact-CSP tie, so the 1.000-vs-2.000 budget reduction and p=9.54e-7 are credible as a mechanism demonstration. The real-repo stress test is honestly negative: the full-pool class is constant by Eq. (7), the candidate-indexed repair is nontrivial but fails its development gate, and Section 8 explicitly states that the obligation-and-risk evidence is a single GLM-5 pass with material earlier evidence-model dependence. This is a real limitation, but the paper's main claim is deliberately a boundary about the tested obligation-and-risk complex; it does not assert a general real-world cohomological advantage. I therefore find no unsupported load-bearing assumption or internal inconsistency that would change the verdict.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a finite 'capability sheaf' model for agent harnesses: typed stalks and literal field-projection restrictions encode local capability signatures, an exact CSP defines global acceptance, and a relative cohomology class over F2 provides a diagnostic score. Four formal results (gluing, portfolio certificate, repair-order separation, rank-truncated spectral stability, and finite-trace recovery) are proved. In a preregistered controlled experiment over 20 task clusters, quotienting hidden interior mediators reduces the first-success candidate budget from 2.000 to 1.000 relative to a stale raw score, and the aligned-state negative control removes the gap. In a real-repository stress test on 160 issues from 20 PatchFuseBench repositories, the full-pool cokernel class is constant by Eq. (7); a candidate-indexed repair is nontrivial on 848/875 candidates but yields only +2 issues over a matched selector (p=0.75) and a LOO abstention ties the anchor (p=1.0), so the development gate fails and the confirmatory split remains sealed. The paper's stated main result is the boundary: controlled invariance holds, while the tested obligation-and-risk complex does not support a real-world cohomological advantage.","tokens_in":15428,"tokens_out":7930,"duration_ms":80121,"significance":"If the results stand, the paper makes a useful methodological contribution: it demonstrates a controlled mechanism by which quotienting hidden-state nuisance variables removes a spurious search gap, and it provides an honest, carefully bounded negative result on a real repository task. The manuscript is unusually disciplined: the controlled design is frozen in advance, exact CSP is used as the semantic control, an aligned-state ablation isolates the mechanism, repository-level inference is used rather than issue-level pseudoreplication, and the confirmatory split is kept sealed after the development gate fails. The mathematical statements are mostly standard but are presented with clean hypotheses and counterexamples, and the proof-backed algebraic identity in Eq. (7) correctly identifies a genuine identifiability defect in the full-pool construction. The negative real-world result is reported with exact p-values and explicit threats to validity, which strengthens rather than weakens the paper. The work is not a general optimizer breakthrough, and the authors do not claim one.","major_comments":[],"minor_comments":[{"comment":"The cochain dimensions dim C0=246, dim C1=233, rank δ0=173, and the derived H1 dimensions are asserted without derivation; please add the incidence-graph counts or point to the artifact script that computes them so a reader can verify the numbers.","section":"Section 2.3"},{"comment":"The sentence 'In all 4,000 registered candidate–target–intervention checks, the quotient signature is unchanged' is presented alongside empirical outcomes; it would be clearer to state that this is a software-verified algebraic identity implied by Eq. (5), not an empirical observation.","section":"Section 6.1"},{"comment":"The abstract writes the budget reduction as '2,000 to 1,000' while the body uses '2.000 to 1.000'; please standardize the decimal notation to avoid confusion with thousands separators.","section":"Abstract and Figure 3"},{"comment":"The identical token rows for exact CSP and full relative class follow by construction because both policies stop at the same first-success candidate; the prose says this, but adding a table footnote would prevent misreading.","section":"Section 5.4 / Table 2"},{"comment":"The 'strong repository-held-out selector' is central to the abstention benchmark but is described only briefly; please add a sentence or two on how it is constructed or cite the artifact location so the anchor comparison is reproducible.","section":"Section 7.1"},{"comment":"The statement that 'Independent GLM-5 and Qwen hunk studies earlier in development showed material evidence-model dependence' is an important threat but has no pointer to where those studies are documented; adding a reference or artifact path would make the limitation auditable.","section":"Section 8"}],"recommendation":"minor_revision","confidential_remarks":"No confidential concerns. The manuscript's honest reporting of the negative real-repository result and the sealed confirmatory split is exemplary, and the minor items above are presentation-level."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the quick version: this paper is worth reading mainly for its discipline and its negative result. It applies cellular sheaf cohomology to agent-harness repair, and it carefully demonstrates a controlled mechanism—quotienting a stale hidden mediator state cuts the candidate budget from 2 to 1 in all 20 clusters, with an aligned-state ablation removing the gap. The real-repository stress test then fails its preregistered gate, and the author says so plainly. That honesty is the best thing about it.\n\nWhat's actually new: the hidden-state quotient construction (Eqs. 4–5), the identifiability counterexample (Eq. 7) showing the naive full-pool class is constant and cannot rank patches, and the candidate-indexed repair (Eq. 8). The controlled experiment is well built: frozen protocol, exact CSP as semantic control, negative control in the aligned condition, and cluster-level statistics. The real test uses 875 real candidates and 153 freshly executed patches, with repository-level p-values and a sealed confirmatory split. That's real evidence.\n\nThe soft spots are real but mostly acknowledged. The quotient result is, by Eq. (5), exactly the public endpoint mismatch, so 'quotient matches exact CSP' is a consistency check on the algebra, not an empirical discovery about the world. The paper says this. The real-repo evidence rows come from a single GLM-5 pass, and the paper admits earlier development showed material evidence-model dependence. So the negative result is about this specific obligation-and-risk complex, not a general statement about cohomological methods. The baseline has only four independent task clusters, and the relative cohomology dimensions in Section 2 are asserted without derivation. Minor.\n\nWho it's for: people working on formal analysis of agent harnesses, sheaf methods in AI, or anyone who wants a clean example of preregistered evaluation and honest negative reporting. The math is standard but well presented.\n\nRecommendation: send it to peer review. A referee should check the artifact and ask for the dimension derivation and more on GLM-5 sensitivity, but the paper is internally consistent and the central claims are bounded. I'd take it.","headline":"Honest boundary paper: controlled invariance mechanism demonstrated, real-repo negative result reported cleanly; worth serious peer review.","tokens_in":15922,"tokens_out":3374,"would_cite":true,"duration_ms":33236,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["18F20","55N30","68T20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish a boundary: quotienting hidden-state nuisance in a capability sheaf halves the candidate search budget in controlled agent-harness tasks, but the same cohomological repair does not beat matched evidence on…","keywords":["capability sheaves","agent harness repair","relative cohomology","constraint satisfaction","hidden-state quotient","patch fusion","gluing obstruction","real-repository evaluation"],"falsifier":"On the same 160-issue development split, rerun the candidate-indexed repair with a second, independently constructed evidence model for edit obligations and risks, or with executed-test feedback replacing model scores; if the method then clears the preregistered gate of at least four issue gains, a positive repository-macro effect, six nonzero repositories, and $p \\leq 0.2$, the paper's claim that the tested obligation-and-risk complex is too weak for real patch fusion would be overturned.","tokens_in":15036,"feed_emoji":"🧩","tokens_out":8422,"duration_ms":84369,"temperature":0.7,"pith_summary":"This paper tries to establish a precise boundary for a sheaf-theoretic repair method for agent harnesses: quotienting out a hidden mediator's stale state reliably halves the candidate search budget in a controlled task family, while the same cohomological construction does not improve real repository patch fusion. The model represents an agent harness as a finite capability sheaf whose stalks are typed behavior signatures and whose restrictions are literal field projections; an exact constraint-satisfaction problem decides acceptance, and a linearized relative cohomology class serves as a diagnostic and search score. In the controlled experiment, all 20 task clusters show the quotient policy reaching first success in one candidate evaluation versus two for the raw stale representative, and aligning the hidden state removes the gap. On the real benchmark, the first pool-level class is constant and cannot rank candidates; a candidate-indexed repair becomes discriminative but resolves only two more issues than a matched noncohomological selector, its abstention gate ties the anchor, and the preregistered development gate fails. The paper's contribution is therefore a boundary: an invariance mechanism that holds in a controlled family and a real-world transfer that fails its own preregistered gate.","feed_headline":"Hidden-state quotient halves search cost, then fails on real repos","feed_subtitle":"A sheaf-theory repair only helps when hidden state is the nuisance; real patch fusion needs stronger evidence.","key_machinery":"The central object is a finite capability sheaf: a cellular sheaf on a five-vertex incidence graph whose vertices are the requirements localization, contract, ordering, preservation, and verification, with typed behavior-signature stalks and restriction maps that are literal field projections. The acceptance rule is the exact CSP of Equation (1): every local signature must lie in its registered good subset and every registered pair of restrictions must agree. The diagnostic is the relative cellular-sheaf cohomology class $\\partial[s_A] = [\\delta^0 s_A] \\in H^1(X,A;\\mathbb{F})$, which vanishes exactly when adjacent typed restrictions agree. The controlled experiment subdivides each overlap with a hidden mediator vertex; changing the hidden value adds an interior coboundary $D_j u = (u,u)$, and the quotient map $P_j = [I\\ I]$ kills it, leaving the public endpoint mismatch $\\ell_j + r_j$. The real experiment's repair indexes the complex by the candidate action, replacing the constant pool-level class with per-candidate scores $q(S)$ computed from candidate-restricted matrices $D_S$. The exact CSP is the semantic control throughout; cohomology is only an invariant diagnostic, never a replacement for exact feasibility or execution.","core_discovery":"The central discovery, stated on the paper's own terms, is that hidden-state quotienting works as a mechanism but not as a real-world advantage. In the controlled family, inserting a hidden mediator on each overlap coordinate makes the raw residual depend on the mediator's value, while the quotient class $[q_j(h_j)]$ maps to the public endpoint mismatch $\\ell_j + r_j$ and is independent of $h_j$; this is what lets the quotient policy stop after one candidate evaluation in every one of the 20 clusters, compared with two for the stale-raw policy, and the aligned-interior ablation confirms that the gain is specifically about removing a nuisance representative. The real-repository stress test then supplies the boundary: because $[b - Dx] = [b]$ in $\\operatorname{coker} D$ for any two selections $x_1$ and $x_2$, a single pool-level class cannot rank candidate configurations. The candidate-indexed repair changes the complex per candidate and is nontrivial on 848 of 875 candidates and varies within 120 of 160 issues, but its direct selection gain is two issues at exact repository sign-flip $p = 0.75$ and its leave-one-repository-out abstention ties the strong anchor at 127 of 160 with $p = 1.0$, so the preregistered development gate fails and the confirmatory split stays sealed. The exact CSP with the same restrictions matches or exceeds the quotient in every controlled setting, so the claim is invariance to stale representatives, not superiority over exact reasoning.","pith_inferences":["Editorial extension: the identifiability defect likely applies to any method that attaches one quotient class to an entire candidate pool rather than to each candidate, so the candidate-indexed repair is a template for other topological or linear diagnostics in decision problems.","Editorial extension: since the controlled gain disappears when the hidden state is aligned, the mechanism predicts that production harnesses with synchronized caches will show no benefit from quotienting; measuring cache-staleness rates in deployed agents would test how often the controlled situation arises in the wild.","Editorial extension: the real-benchmark evidence model is the obvious lever—replacing the single model pass with executed-test feedback or a second independently generated evidence pass on the same 160 issues would tell whether the boundary is fundamental or just a limitation of the available semantic rows."],"forward_implications":["If the controlled invariance result is correct, agent-harness searches that score raw half-edge residuals can be misled by stale cached representatives, and quotienting those interior coboundaries removes the artifact without changing the public gluing problem.","If the identifiability counterexample is taken at face value, a single pool-level cohomology class can never rank candidate configurations, so any future linear diagnostic must index the complex by the candidate or action under consideration.","If the real-repository negative result is accepted as the present boundary, cohomological patch fusion will not beat matched semantic evidence until the obligation-and-risk rows capture actual edit-edit semantic conflicts rather than model-generated judgments.","In every tested setting, the exact CSP with the same restrictions matches or exceeds the quotient, so cohomological search scores are complements to, not replacements for, exact feasibility and execution."],"supporting_citations":[{"why":"supplies the fixed real-repository candidate pool and the deterministic patch-fusion baseline that the stress test reuses and compares against.","marker":"(Yang et al., 2026)"},{"why":"supplies the real GitHub-issue pool from which the development and confirmatory splits are drawn.","marker":"(Jimenez et al., 2024)"},{"why":"motivates sheaf-theoretic local-consistency witnesses for failures that lack global sections.","marker":"(Abramsky and Brandenburger, 2011)"},{"why":"supplies the cellular-sheaf and relative-cohomology formalism used for the finite complex.","marker":"(Curry, 2014)"},{"why":"supplies the singular-subspace perturbation bound used in the rank-truncated stability theorem.","marker":"(Wedin, 1972)"},{"why":"supplies the concentration inequalities used in the finite-trace recovery and stability guarantees.","marker":"(Hoeffding, 1963; Azuma, 1967)"}],"fun_headline_variants":["Hidden-state quotient halves search, but real repos see no advantage","Sheaf quotient fixes hidden nuisances, fails on real patch selection","Cohomology repair: controlled gain, real-world tie at best","Quotient invariance helps, but exact CSP matches; repo test seals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The real-repository negative result rests on a single pass of model-generated obligation-and-risk evidence over edit atoms, and if that evidence misses the true semantic conflicts between edits, the failure belongs to the evidence model rather than to the cohomological method.","fun_headline_variants_meta":{"raw":{"variants":["Hidden-state quotient halves search, but real repos see no advantage","Sheaf quotient fixes hidden nuisances, fails on real patch selection","Cohomology repair: controlled gain, real-world tie at best","Quotient invariance helps, but exact CSP matches; repo test seals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2946,"prompt_tokens":1232,"completion_tokens":1714,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":848,"completion_tokens_details":{"reasoning_tokens":1637}},"tokens_in":848,"tokens_out":1714,"duration_ms":13869,"temperature":1.0,"reasoning_tokens":1637,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:48:30.100535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the same 160-issue development split, rerun the candidate-indexed repair with a second, independently constructed evidence model for edit obligations and risks, or with executed-test feedback replacing model scores; if the method then clears the preregistered gate of at least four issue gains, a positive repository-macro effect, six nonzero repositories, and $p \\leq 0.2$, the paper's claim that the tested obligation-and-risk complex is too weak for real patch fusion would be overturned.","supporting_citations":[{"cited_title":"A Single Patch Is Not Enough: Deterministic Fusion of Repair Candidates , howpublished =","cited_arxiv_id":null,"evidence_quote":"supplies the fixed real-repository candidate pool and the deterministic patch-fusion baseline that the stress test reuses and compares against."},{"cited_title":"2014 , note =","cited_arxiv_id":null,"evidence_quote":"supplies the cellular-sheaf and relative-cohomology formalism used for the finite complex."}],"review_version":1}