{"id":"d52bfdf7-25e1-4f55-afa1-d9414c0968b2","arxiv_id":"2607.03695","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Attention width and source social power determine whether LLM agent networks herd or achieve wisdom-of-crowds, with a pricing equalizer restoring optimal collective weights.","lead":"LLM agent populations can herd into false consensus or pool real knowledge depending on how sharply each agent focuses attention. The paper gives a control knob and a pricing fix that move groups between those two regimes.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Bridge/anchoring mismatch is real but secondary; the load-bearing gap is that primary accuracy claims rest on mechanistic analogy rather than non-vacuous transfer of Neff.","rationale":"The reader correctly isolates the bridge/anchoring mismatch as the weakest structural assumption. I agree it is the single most load-bearing concern for the packaged strongest claim, because the theorems that give the clean Neff≤8 / Neff=n statements do not non-vacuously cover the primary accuracy figures. The paper is transparent about this (§5.2, Prop. 8.4) and supplies independent support: controlled testbed with λ_max<1, large ablated accuracy swings, equalizer recovery, placement effect, and cross-family replication. Those keep the contribution accept-shaped; the gap is that “reproduces” is currently mechanistic analogy plus testbed certification rather than theorem-to-plot transfer. The concrete re-run under anchoring would settle whether the accuracy claims inherit the proved Neff bounds or remain correlational. No stronger internal inconsistency appears; the math is self-contained and the empirical operator effects are large. Verdict stays CONDITIONAL; confidence remains moderate for the same reasons the reader gave.","tokens_in":40025,"tokens_out":752,"duration_ms":6843,"concrete_test":"On the same HiddenBench/Werewolf seeds and hub adversary, re-run the full β grid with the explicit anchoring-persona prompt of App. 12.7 (so fitted λ_max stays comfortably below 1) and report both collective accuracy and the measured Neff(q_T) / bridge residual δ_T/(1-λ_max) side-by-side with the unanchored curves. If the accuracy swing collapses or Neff no longer tracks accuracy once the bridge is non-vacuous, the primary claim weakens; if both the swing and the Neff–accuracy correspondence survive, the concern is largely aesthetic.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim packages Theorems 4.2–4.3 (Neff≤8 under dominant pair; Neff\to n under double stochasticity) with the statement that the herding–wisdom transition “reproduces” on HiddenBench/Werewolf. Those theorems characterize the proxy weight q_T (or its stationary \nu_C) under the realized-influence operator C(eta). Theorem 4.1 transfers proxy moments to the LLM only under δ-approximate anchored best response with λ_max<1. The primary accuracy plots (Fig. 2A,B; Tables 5–9) use DeGroot updating (λ=1), where Proposition 8.4 shows the bridge gap can grow linearly and the bound is vacuous. The authors state this explicitly (§5.2): the discussion benchmarks share the β-gating mechanism, “not a theorem mapping Neff onto accuracy.” Anchored-testbed certification (Fig. 4, 254 cells) and OLS slope 0.88 are genuine, but they live on a different regime (median fitted λ=0.25) from the headline plots. Consequently the central empirical claim—that operator-controlled β swings accuracy because Neff collapses/recovers—is supported by large, ablated accuracy swings and a shared attention bottleneck, not by a non-vacuous path-wise link from the proved Neff bounds to those accuracy numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces SNLA, a social-network model for populations of LLM agents that separates visible exposure from realized influence. Realized influence is derived from exposure weights, Katz–Bonacich social power, and a finite attention width β via a tempered (temperature-softmax) allocation. On a tractable Friedkin–Johnsen proxy, the authors prove a path-wise bridge to δ-approximate LLM emissions under anchoring (Theorem 4.1), a two-regime result in which narrow β yields herding with Neff bounded independently of n while wide β recovers wisdom-of-crowds accuracy only for undirected degree-regular exposure (Theorem 4.2), and an equalization theorem showing that a decentralized Sinkhorn-style pricing protocol drives collective weights to uniform when column defects vanish (Theorem 4.3). Empirically, a controlled scalar testbed validates Neff/variance predictions and the bridge; operator-controlled variants of HiddenBench, Werewolf, and AgentsNet show large accuracy/coordination swings with β, placement effects, and equalizer recovery of herding.","tokens_in":40310,"tokens_out":1390,"duration_ms":19214,"significance":"If the results hold, the paper supplies a usable control (attention width β, plus an equalizer) for when multi-agent LLM systems pool information versus herd—directly relevant to debate, simulation, and collaborative agent systems. Strengths include explicit, self-contained proofs with stated assumptions; a measurable residual δ that makes the bridge falsifiable; a controlled testbed that reports Neff and the bridge bound (254/254 cells); large, ablated accuracy swings on recognized benchmarks; and cross-family replication on Llama. The MoE-routing analogy and the pricing protocol as information projection are additional conceptual contributions. The work is a genuine advance over classical DeGroot-style models that treat exposure as influence.","major_comments":[{"comment":"§5.2 and Theorem 4.1 / Proposition 8.4: The primary collective-accuracy claims (Fig. 2A,B; Tables 5–9) use DeGroot updating (λ=1), where the authors correctly note the bridge bound is vacuous and that the discussion benchmarks share the β-gating mechanism rather than a theorem mapping Neff onto accuracy. The non-vacuous bridge evidence (Fig. 4; OLS slope 0.88, R²=0.74) lives on the anchored testbed (median fitted λ≈0.25). This is load-bearing for the packaged claim that the herding–wisdom transition “reproduces” the theory: please either (i) report proxy Neff or q_T diagnostics on the discussion runs themselves, or (ii) reframe the headline empirical claim more sharply as a mechanistic demonstration of the attention bottleneck, with theorem-backed Neff transfer reserved for the anchored testbed.","section":"§5.2, Theorem 4.1, Fig. 2A–B"},{"comment":"Appendix 7 / §5.2 (bridge certification): Anchoring weights λ_i are least-squares fit from the same residuals ˆδ_i(t) that enter the uniform residual δ_T used in the bridge bound. After per-agent best-fit λ, “0 violations on 254 cells” is a statement about best-fit ceilings, not about a prescribed anchoring regime. The manuscript already notes this; please make the distinction fully explicit in the main text (not only the appendix) and report the distribution of fitted λ and of δ_T/(1−λ_max) separately for anchored-persona vs. unanchored cells so readers can judge how non-vacuous the bound is on the subset that actually supports Theorem 4.1.","section":"Appendix 7, §5.2, Fig. 4"},{"comment":"Theorem 4.2 (narrow regime): The Neff≤8 bound relies on a unique dominant pair (j★_i = h for all i≠h, j★_h = g≠h) and ρ(β)<1. This is a strong structural assumption. Please state how often the unique-dominant-pair condition holds on the exposure graphs of HiddenBench/Werewolf (or give a weaker multi-source concentration bound), and whether the empirical herding floor is consistent with concentration on a small set rather than specifically a pair. Without this, the quantitative Neff≤8 prediction is only loosely connected to the benchmark herding floors.","section":"Theorem 4.2, §9.3"}],"minor_comments":[{"comment":"Figure 1 is dense; the herding vs. wisdom cartoons and the operator bridge formula compete for space. Consider splitting the schematic from the task vignette.","section":"Figure 1"},{"comment":"Notation: π is social power and ν is the consensus weight of C; both are stationary-like objects. A one-line reminder at first use of ν_C(β) would help.","section":"§4.3, Theorem 4.2"},{"comment":"Table 1 and Appendix 13.2: the equalizer can slightly raise wide-β error in some cells (e.g., k=4). A brief sentence on when equalization can overshoot an already-balanced allocation would prevent misreading.","section":"Table 1, §5.3"},{"comment":"Boundary environments (Debate, GovSim, MARBLE) are useful; the negative MARBLE placement result is honest. Consider moving the full boundary table into the main text or a short dedicated subsection so the scope of the theory is clearer.","section":"Appendix 13.8"},{"comment":"Minor typos: “APREPRINT” headers; occasional spacing in math (e.g., N_eff formatting). Standard copy-edit pass.","section":"Throughout"}],"recommendation":"minor_revision","confidential_remarks":"The theory–empirics gap on λ=1 is real but the authors already flag it; I would not reject on that basis. The contribution is above the bar for a solid ML venue if the framing of “reproduces” is tightened. Fit is good for multi-agent / collective-intelligence tracks; less so for pure theory venues that demand non-vacuous transfer on every empirical plot."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is a single harness knob—attention width β—that moves multi-agent LLM populations between herding and pooling, plus a cheap equalizer that undoes the collapse. That is not just another debate paper; they separate visible exposure from realized influence via social power and finite attention, then prove Neff bounds on a tractable proxy and show the same swing on HiddenBench and Werewolf.\n\nWhat is new is the operator C built from exposure × power through a temperature-β softmax, the two-regime theorem (narrow β caps Neff independently of n under a dominant pair; wide β recovers the degree ceiling only on undirected regular graphs), and the Sinkhorn-style exposure prices that drive column defects to zero. Classical DeGroot/Friedkin–Johnsen and Golub–Jackson are reused correctly as special cases; the equalizer is one Sinkhorn step per round, which they own. The controlled testbed checks variance and Neff predictions directly; the benchmark swings are large (+0.6 accuracy), gated by ζ, and survive adversary-on/off and model-family checks. Math is self-contained with explicit constants; ablations are honest about incomplete online recovery on short-horizon Werewolf.\n\nThe soft spot the stress-test flags is real but secondary. Theorem 4.1 needs λ_max < 1 and δ-approximate anchored best response; the headline accuracy plots run pure DeGroot (λ = 1), where the authors themselves say the bound is vacuous and they are sharing a mechanism, not a theorem mapping Neff onto accuracy. Anchored-testbed certification (254 cells, slope 0.88) is genuine and lives in a different regime. So the central empirical claim rests on large ablated accuracy swings plus a shared attention bottleneck, not path-wise transfer of the proved Neff numbers. That is a gap in the packaging, not a hole in the operator results. Free parameters (β, ζ, k, coverage) are declared; no public code hash is a practical nuisance, not a soundness issue.\n\nThis is for people building or auditing multi-agent LLM systems who need a control and a fix, and for anyone who cares about opinion dynamics under finite attention. It deserves a serious referee. I would engage with it and expect revision on the bridge language, not rejection.","headline":"Clean operator-level control of herding vs. wisdom in multi-agent LLMs, with real theorems and large ablated accuracy swings; the bridge-to-proxy is only non-vacuous on the anchored testbed, not on the headline DeGroot plots.","tokens_in":41016,"tokens_out":581,"would_cite":true,"duration_ms":6486,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A single attention-width knob decides whether LLM agent populations pool knowledge or herd into false consensus.","keywords":["LLM multi-agent systems","social learning","opinion dynamics","attention width","herding","wisdom of crowds","realized influence","equalized allocation"],"falsifier":"On an aggregation-dependent multi-agent task with a fixed dominant wrong source, measure collective accuracy while sweeping attention width: if accuracy stays flat instead of rising sharply from a low floor at narrow width to near-oracle levels at wide width, or if equalizing column sums fails to lift the narrow-width floor, the central claim fails.","tokens_in":40820,"feed_emoji":"🧠","tokens_out":956,"duration_ms":7885,"temperature":0.7,"pith_summary":"When large language model agents talk to one another, what the group ends up believing can be either wiser than any member or confidently wrong. Classical social-network theory treats the visible links among agents as the whole story of how beliefs combine. That picture fails for language-model agents: each one has a limited context window and selective attention, so it only realizes part of what it is exposed to. The paper introduces SNLA, a framework that separates visible exposure from realized influence. Realized influence is built from each source’s social standing and a single temperature parameter that controls how sharply attention focuses. On a tractable mathematical proxy the authors prove that narrow attention collapses the effective sample size of the group to a constant independent of population size (herding), while wide attention recovers wisdom-of-crowds accuracy only when the exposure graph is balanced. A simple decentralized pricing rule that equalizes attention across sources restores the optimal collective weights. Controlled experiments and three multi-agent benchmarks show the same herding-to-wisdom transition, and the pricing fix removes herding without retraining the models.","feed_headline":"One attention knob flips LLM groups from herding to wisdom","feed_subtitle":"Narrow focus collapses group accuracy to a few sources; wide focus plus balance restores the crowd.","key_machinery":"The realized-influence operator C: each reader reweights its exposure by source social power and then applies a temperature-β softmax, so that C, not the visible network, drives belief updates. A path-wise bridge theorem couples real LLM emissions to an anchored Friedkin–Johnsen proxy, after which two-regime and equalization theorems characterize effective sample size as a function of β and column balance.","core_discovery":"The paper establishes that collective accuracy of an LLM-agent population is controlled by attention width: narrow attention produces herding whose effective sample size stays bounded no matter how large the population grows, while wide attention recovers wisdom-of-crowds behavior only on undirected degree-regular exposure graphs; a decentralized equalizer that drives the influence matrix toward double stochasticity restores optimal collective weights whenever a dominant source is present.","pith_inferences":["The same column-collapse pathology appears in mixture-of-experts routing; the paper’s equalizer is formally the same family of load-balancing fixes, suggesting a shared fix across agent societies and expert routing.","If the bridge can be made non-vacuous for unanchored DeGroot-style debate, the theory would directly certify the accuracy curves already plotted on the main benchmarks rather than only on the anchored testbed.","Operator-controlled variants of existing multi-agent suites could become a standard stress test for whether a new agent architecture herds or pools."],"forward_implications":["Designers of multi-agent LLM systems can treat attention width as a single control that moves a population between herding and wisdom without changing model weights.","When a high-power source is present, one Sinkhorn-style price update per round is enough to keep collective variance near the optimal 1/n floor.","Placement of capable agents at high-degree nodes improves coordination even when there is no single wrong answer to herd onto.","Classical wisdom-of-crowds guarantees transfer to LLM societies only after exposure is made doubly stochastic and attention is sufficiently wide."],"fun_headline_variants":["Narrow attention locks LLM agents into herding at any group size","Wide attention restores LLM wisdom only on balanced undirected graphs","Attention width decides if LLM populations herd or pool knowledge","LLM agents herd under narrow focus; wide balanced focus enables crowds","SNLA shows attention focus bounds LLM group effective sample size"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The mathematical link from real language-model replies to the analyzable proxy requires every reply to stay close to an anchored update whose self-weight is strictly less than one; the main discussion benchmarks run without that anchor.","fun_headline_variants_meta":{"raw":{"variants":["Narrow attention locks LLM agents into herding at any group size","Wide attention restores LLM wisdom only on balanced undirected graphs","Attention width decides if LLM populations herd or pool knowledge","LLM agents herd under narrow focus; wide balanced focus enables crowds","SNLA shows attention focus bounds LLM group effective sample size"]},"model":"grok-4.5","effort":"low","cost_usd":0.004354,"raw_usage":{"total_tokens":1293,"prompt_tokens":755,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":43540000,"prompt_tokens_details":{"text_tokens":755,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":454,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":755,"tokens_out":84,"duration_ms":3610,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T00:36:39.283554+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On an aggregation-dependent multi-agent task with a fixed dominant wrong source, measure collective accuracy while sweeping attention width: if accuracy stays flat instead of rising sharply from a low floor at narrow width to near-oracle levels at wide width, or if equalizing column sums fails to lift the narrow-width floor, the central claim fails.","supporting_citations":[],"review_version":1}