{"id":"af0fb1bd-3e18-4201-a748-15db2b27dabb","arxiv_id":"2608.11357","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across controlled multi-agent experiments, institutional interventions beat more capable models only when they repair the construction of usable public state, and lose that advantage when signals are uncheckable, capability substitutes, or the state is unexecutable.","lead":"This paper tests when adding institutional structure, such as routing, validation, and audit, helps groups of AI agents more than simply using a stronger model. It finds institutions help when they fix how the group builds shared knowledge, but stop helping when the model can do the same work itself or when the shared state cannot be acted on.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Frontier substitution boundary is under-identified: zero valid-minus-broken interaction in Fig. 8a is read as 'capable model substitutes,' but arm means are not reported, so it could equally be insensitivity or null.","rationale":"The reader's confound concern is real, but it targets the positive mechanism evidence; the missing arm means in the frontier panel target one of the headline boundary conditions directly. The central claim is an 'institutions help when... but lose when...' claim; the second clause's strongest direct evidence is the frontier substitution result. Without absolute accuracies, the paper cannot rule out that zero interaction means the model does not use the institutional state at all, which would not demonstrate substitution. The paper elsewhere reports arm means when absolute failure is informative, so the omission in the frontier panel is conspicuous. Since the paper is otherwise careful with mechanism-breaking controls and honest nulls, this is a verification gap rather than a reason to reject; the conditional verdict remains appropriate.","tokens_in":12019,"tokens_out":13761,"duration_ms":129952,"concrete_test":"Report the arm-level accuracies for the frontier panel: for each ecology in Fig. 8a (dynamic state, expertise routing, strategic reporting), list the absolute accuracy of DeepSeek, Gemini, and Claude in the valid and broken conditions, alongside the institution-supported 8B accuracy. If a zero-interaction model is at or above the 8B institutional accuracy in both arms, substitution is confirmed; if it is near chance or clearly below the 8B institution in both arms, the zero interaction is a null or insensitivity, and the frontier-substitution label in Table 1 should be revised to the uninformative-signal condition or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 7, the frontier panel reports only the valid-minus-broken interaction for DeepSeek, Gemini, and Claude (Fig. 8a), and Table 1 labels Claude's 0.00 interaction as 'Frontier substitution.' A zero interaction is compatible with at least three distinct states: (i) substitution, where the raw model achieves high accuracy in both the valid and broken conditions by reconstructing the transformation from raw context; (ii) insensitivity, where the model ignores the institutional signal and performs at the same low accuracy in both arms; and (iii) a floor or ceiling null. The text asserts interpretation (i) ('when a capable reasoner can reliably perform the same transparent transformation from raw context'), but the reported statistic does not distinguish it from (ii) or (iii). This matters because the abstract's third boundary condition—institutions 'lose their advantage when stronger intelligence can perform the same transformation directly'—rests on this panel. If Claude's zero interaction is actually low-and-equal accuracy, the result belongs under the 'uninformative signal' boundary, not the capability-substitution boundary, and the central claim's 'relative to a capability frontier' clause lacks direct model-level support. The learned-reranker comparison on MuSiQue provides substitution evidence for evidence construction, but not for the dynamic-state, expertise-routing, or reporting transformations tested in the frontier panel.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks under what conditions institutional design outperforms additional model capability in multi-agent LLM collectives. It formalizes an institution as externally specified rules for observation, admission, state updating, and action interface, and constructs four synthetic ecologies around access/routing, admission/dependence, state maintenance/incentives, and representation/action. Positive institutional interventions are paired with call-matched raw baselines and mechanism-breaking controls, with additional tests on HotpotQA, MuSiQue, and frontier models. The central claim is that institutions help when they repair failures in constructing usable public state, but lose their advantage when their signals are uninformative or uncheckable, when stronger capability can perform the same transformation directly, or when the resulting state cannot support reliable action. The paper is explicitly framed as a diagnosis of where collective reasoning fails rather than as a universal ranking of workflows.","tokens_in":12311,"tokens_out":11263,"duration_ms":95063,"significance":"If the results hold, the paper offers a valuable empirical methodology for deciding when to invest in model capability versus institutional structure in multi-agent systems. Its strengths include controlled manipulations with paired baselines, mechanism-breaking controls (broken routing, syntax-only validation, zero checkability, fixed-evidence interfaces), honest reporting of nulls and wide intervals, and transfer to natural tasks (HotpotQA, MuSiQue). The manuscript is falsifiable in spirit, and it explicitly does not claim that institutions are universally superior. The taxonomy of collective failure loci, and the insistence on separating construction of public state from execution on it, are useful contributions regardless of the specific point estimates.","major_comments":[{"comment":"The zero valid-minus-broken interaction for Claude in all three frontier ecologies is interpreted as evidence that a capable reasoner can reliably perform the same transparent transformation from raw context. With only the interaction reported, this zero is compatible with at least three states: high accuracy in both arms (substitution), low accuracy in both arms (the model ignores the institutional signal), and a floor or ceiling null. The abstract's third boundary condition, that institutions lose their advantage when stronger intelligence can perform the transformation directly, rests on this panel, so the arm means for the valid and broken conditions for each frontier model must be reported, or an accuracy measure for the direct transformation must be provided. If Claude's zero interaction is actually low-and-equal, the result belongs under the 'uninformative signal' boundary rather than a capability-substitution boundary. The MuSiQue learned-reranker comparison supports substitution only for evidence construction, not for the dynamic-state, expertise-routing, or reporting transformations tested in Fig. 8a.","section":"Section 7, Fig. 8a, Table 1"},{"comment":"The 3x2 and 3x4 routing conditions vary the number of clue cards exposed (6 and 12 cards, respectively) alongside routing completeness, so the reported chain 'routing completeness -> decisive-evidence coverage -> unique public decision state -> final accuracy' is confounded with information quantity. The claim that below-threshold routing tests 'whether coverage rather than workflow complexity matters' is not cleanly supported unless card count is held fixed across routing conditions. The paper should state explicitly whether the raw baselines in the headline +0.58 comparison also receive all 16 clue cards under matched calls; if they do not, the institutional gain is not separated from simply showing more evidence to the finalizer.","section":"Section 4, Fig. 4 and below-threshold routing"},{"comment":"The formal definition of the mechanism-validity interaction is not fully specified. The notation I^r_0 is never defined, and as written tau^r_I(m) is a difference relative to an unspecified baseline, so Gamma_I(m) is not transparently equal to the valid-minus-broken interaction used in the experiments. Because this definition is introduced as the justification for giving mechanistic credit only when performance tracks the institutional signal, please define the baseline institution explicitly or remove the formal equations in favor of the verbal definition used in the experimental sections.","section":"Section 2, Eq. (1)"}],"minor_comments":[{"comment":"In the sentence 'the same passages are presented either through the graph-style public state or through an ordered or passage-pointer interface,' it is unclear whether the ordered interface and the passage-pointer interface are the same condition or two different conditions; please clarify and state whether finalizer call counts are matched in this fixed-evidence comparison.","section":"Section 7, fixed-evidence interface"},{"comment":"The caption says points are arm accuracies with n=50 and that the bottom row reports the paired valid-institution minus stronger-raw effect; adding the call count to each labeled method row (e.g., 5-call 8B vs 5-call 70B) would make the matched-calls claim easier to verify from the figure alone.","section":"Figure 4 caption"},{"comment":"The limitation paragraph concedes that call matching does not equal token, latency, monetary cost, or engineering burden; this caveat is important framing and should also appear where 'matched reasoning baselines' is first introduced in Section 1, to avoid leaving the impression that the comparison is a matched-resource comparison.","section":"Section 1 and Section 10, matched calls"},{"comment":"HotpotQA uses 20 paired questions per model; given the small sample, the paper should state this sample size explicitly in the HotpotQA subsection and connect it to the reported bootstrap intervals, since the graph-routing answer effects may be wide.","section":"Section 3 and Section 7, HotpotQA sample size"},{"comment":"The names 'Qwen 235B', 'Mistral 24B', and 'Claude Sonnet 4.6' appear without version qualifiers; please give the exact model identifiers used in the API calls so that the results are reproducible.","section":"Section 3, model naming"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope and the experimental design is thoughtful. The main technical concern is the under-identified frontier substitution panel in Fig. 8a; if the authors can report arm means or a direct transformation accuracy measure, I would be willing to reconsider. A reproducibility appendix with exact prompts, card decks, and seed sets would also strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a thoughtful empirical paper that mostly delivers on its central claim, but the headline frontier-substitution boundary is under-identified and the lack of artifacts makes verification harder than it should be.\n\nWhat's genuinely new: the matched mechanism-breaking methodology. The paper doesn't just show that adding roles or voting helps; it pairs each institutional intervention with a call-matched stronger raw model and with a broken-signal control, so gains are credited only when performance tracks the signal the institution claims to exploit. The four-locus taxonomy (access/routing, admission/dependence, maintenance/incentives, representation/action) is useful, and the paper is honest about nulls: dynamic-state benefits are weak, the hosted strategic pilot found no misreporting, and graph answer transfer is imprecise. The crossover results in the synthetic ecologies are large and internally consistent.\n\nThe main soft spot is exactly what the stress-test note says. The frontier panel (Fig. 8a) reports only valid-minus-broken interactions for DeepSeek, Gemini, and Claude. A zero interaction is compatible with substitution, insensitivity, or floor/ceiling. The paper reads it as substitution without reporting arm means. That matters because the third boundary condition in the abstract — institutions lose their advantage when stronger intelligence can perform the transformation directly — rests on this panel. The learned-reranker result on MuSiQue supports substitution for evidence construction, but not for the dynamic-state, expertise-routing, or reporting transformations. So the capability-frontier clause is plausible but not directly established for those ecologies. That's a moderate flaw, not fatal: the overall boundary claim survives on the other controls.\n\nTwo lesser issues. The below-threshold routing (3x2, 3x4) varies the number of clue cards alongside routing completeness, so the coverage story is entangled with information quantity. The paper acknowledges call matching doesn't equal token/cost matching, but it's still a limit on the stronger-than claim. And there's no code or data release, which for an empirical paper of this kind is a real omission — the reader shouldn't have to trust the bootstrap intervals on faith.\n\nFor whom: anyone working on multi-agent LLM systems, institutional design, or collective intelligence. It's a serious paper, worth a full review. I'd send it to referees and ask for arm means in the frontier panel, a re-analysis of the routing manipulation holding card count fixed if possible, and artifact release. That's the right scope of revision.","headline":"A serious, mostly sound empirical boundary on when institutions beat scaling in multi-agent LLMs; the frontier-substitution panel needs arm means before the boldest claim is fully supported.","tokens_in":12783,"tokens_out":3008,"would_cite":true,"duration_ms":83283,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that institutional rules for collective state construction beat additional agent capability only when they repair a structural failure in routing, evidence admission, checkable enforcement, or action representation, and…","keywords":["multi-agent systems","institutions","collective intelligence","group decision making","large language model agents","empirical methodology","distributed cognition","checkability"],"falsifier":"Give the stronger raw reasoner on the distributed-evidence ecology the same token budget, latency, or dollar cost as the institutional pipeline instead of only the same call count, and measure whether the +0.58 paired advantage persists; if the advantage collapses while routing completeness remains unchanged, the claim that institutions beat intelligence by repairing coverage rather than by spending more computation is undercut.","tokens_in":11822,"feed_emoji":"🏛️","tokens_out":9682,"duration_ms":78805,"temperature":0.7,"pith_summary":"More capable agents do not automatically form a more capable collective, and this paper tries to pinpoint when the fix should be a stronger reasoner and when it should be a change in the rules through which agents pool information. Using controlled artificial ecologies built around four failure loci—access and routing, admission and dependence, state maintenance and incentives, and representation and action—it finds a consistent boundary: institutions help when they repair how the collective constructs a usable public state. The advantage disappears when the institutional signal is uninformative or uncheckable, when stronger intelligence can perform the same transformation directly, or when the constructed state cannot support reliable action. The upshot is that choosing between intelligence and institutions should be a diagnosis of where collective reasoning fails, not a general preference for one over the other.","feed_headline":"Institutions beat raw intelligence only when public state is broken","feed_subtitle":"Rules win only when they repair the path from private evidence to public action; stronger models erase the edge.","key_machinery":"The carrying object is the formal institution as state-construction rule: $I = \\langle \\rho, V, U, \\Phi \\rangle$, where $\\rho$ is who may observe, report, or route; $V$ is which reports are admitted as public evidence; $U$ is how admitted evidence updates shared state; and $\\Phi$ is how that state is represented to the finalizer. The paper pairs this with two controlled comparisons: a mechanism-validity interaction $\\Gamma_I(m) = \\tau_I^+(m) - \\tau_I^-(m)$, which asks whether performance tracks the informational signal the institution claims to use, and a crossover $C_I(m_s, m_\\ell)$ that compares an institutionally supported weaker reasoner with a stronger raw reasoner under matched calls. These definitions, applied inside ecologies built around access, admission, maintenance, and representation, turn \"institutions versus intelligence\" into a test of whether the proposed rule is informative, checkable, substitutable, and executable.","core_discovery":"The paper's central discovery is that whether institutions beat intelligence is an empirical, locatable fact rather than a stylistic preference. Across four controlled ecologies, an institution—a fixed rule for who may report, what is admitted into public belief, how shared state is updated, and how it is exposed to the final decision—improves collective accuracy only when it repairs a specific failure along the path from private observation to public action. Concretely: validated, coverage-complete routing lets a five-call 8B model beat a five-call raw 70B model by +0.58 paired accuracy on a distributed-evidence task; grounded admission takes faulty-evidence accuracy from 0.00 to 1.00; a 10% checkable audit at capability 0.55 beats an unaudited capability 0.65 by +0.135, while all four sanction levels at zero checkability produce exactly zero gain; and a learned reranker outperforms a hand-built graph, while a fixed-evidence interface reversal shows that better evidence does not guarantee better execution. The same boundary conditions define the limits: uninformative or uncheckable signals, capable raw reasoners that can perform the transformation directly, and unexecutable representations all erase the institutional advantage.","pith_inferences":["A natural extension the paper does not run: an adaptive controller that measures which failure locus is alive (routing coverage, admission validity, checkability, interface executability) and switches between raw reasoning and institutional machinery on that diagnosis; the paper's Gamma and crossover comparisons are exactly the hooks such a controller would need.","Because the paper's resource accounting stops at call counts, a reader should expect institutional advantages to shrink when compared at equal token, latency, or dollar budgets; if they do, the boundary would be economic rather than purely cognitive.","The zero-checkability exact-zero result suggests a testable design rule for LLM workflows: any governance mechanism whose outputs cannot be independently checked by downstream agents should be predicted to contribute zero marginal accuracy, which could be probed by varying only the verifiability of a reviewer's output while holding its content distribution fixed."],"forward_implications":["In a workflow whose failure is poor evidence routing, buying a larger or more powerful model is the wrong intervention: the distributed-evidence results show complete routing plus validation converts a 0.26-accuracy 3x2 assignment into 1.00 under the disjoint 4x4 assignment, whereas call-matched raw strength on the same partial state stays near the floor.","An institutional rule earns causal credit only when performance tracks its signal: syntax-only validation, broken credentials, and uncheckable sanctions produce no gain, so adding such machinery is pure overhead.","Institutional advantage is capability-relative: as models or learned retrieval components become able to perform the same transformation from raw context, the permanent institution becomes unnecessary, as shown by the frontier substitution and learned-reranker-over-graph results.","Construction and representation must be evaluated separately: improving the public state's evidence quality does not guarantee better final decisions, as shown by the fixed-evidence interface reversal where the ordered interface flips the graph comparison.","The practical rule of thumb the paper supports: buy intelligence when capability can perform the transformation; build an institution when an active structural failure survives scaling, the rule is checkable, and the resulting state is executable; otherwise redesign the interface."],"supporting_citations":[{"why":"Meta-analysis of hidden-profile experiments; supplies the access-and-routing failure mode in which groups possess but do not pool unshared evidence.","marker":"[17]"},{"why":"Shows how knowledge about the distribution of expertise improves coordination; motivates the expertise-routing and coverage-routing institutions.","marker":"[25]"},{"why":"Establishes that repeated agreement can reflect social influence rather than independent evidence; motivates the admission-and-dependence ecology.","marker":"[1, 2, 12]"},{"why":"Accountability research showing that consequences depend on observability; motivates the checkability condition for sanctions.","marker":"[23, 27, 28]"},{"why":"Distributed-cognition and external-representation work; motivates the representation-and-action locus and the fixed-evidence interface tests.","marker":"[15, 22, 34]"},{"why":"Institutional analysis and normative multi-agent systems; supplies the definition of institutions as rules governing collective state construction.","marker":"[19, 20, 30]"},{"why":"Explainable multihop QA dataset; used to test whether a title-link graph institution constructs better public evidence than equal-size lexical selection.","marker":"[33]"},{"why":"Harder multihop question dataset; used to compare hand-specified graph construction with sparse and dense retrieval plus a learned reranker, probing substitution.","marker":"[29]"}],"fun_headline_variants":["Institutions only repair broken public state","Institutional edge fades when public state works","Rules matter only when public state is broken","Intelligence beats institutions when state is sound"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The attribution of gains to institutional repair rather than to extra compute or extra information assumes that matching the number of reasoning calls, together with mechanism-breaking controls, isolates the institutional mechanism; the paper itself notes that call matching is not the same as matching tokens, latency, or monetary cost.","fun_headline_variants_meta":{"raw":{"variants":["Institutions only repair broken public state","Institutional edge fades when public state works","Rules matter only when public state is broken","Intelligence beats institutions when state is sound"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001108,"raw_usage":{"total_tokens":4647,"prompt_tokens":1000,"completion_tokens":3647,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":3592}},"tokens_in":616,"tokens_out":3647,"duration_ms":25710,"temperature":1.0,"reasoning_tokens":3592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:13:12.527567+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the stronger raw reasoner on the distributed-evidence ecology the same token budget, latency, or dollar cost as the institutional pipeline instead of only the same call count, and measure whether the +0.58 paired advantage persists; if the advantage collapses while routing completeness remains unchanged, the claim that institutions beat intelligence by repairing coverage rather than by spending more computation is undercut.","supporting_citations":[{"cited_title":"Vaughan, and Dennis D","cited_arxiv_id":null,"evidence_quote":"Shows how knowledge about the distribution of expertise improves coordination; motivates the expertise-routing and coverage-routing institutions."}],"review_version":1}