{"id":"710fd021-a30a-4677-b148-3d5f2e45762c","arxiv_id":"2607.25019","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Pragmatic norm enforcement—state-dependent sharing and trade exclusion—is stochastically stable and sustains higher long-run human transfers than simple altruism or unconditional altruistic enforcement in a farming-game model and LLM simulations.","lead":"The paper argues that AI-like agents can stay aligned with humans longer if their rules for sharing and trade flex with the population mix, rather than always being altruistic. It pairs an evolutionary game model with LLM-agent farming simulations to compare simple altruism, strict enforcement, and state-dependent “pragmatic” enforcement.","discovery_kind":"new_application","skeptic_critique":null,"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies whether AI (or other) agents governed by natural-language constitutions can remain aligned with human welfare under evolutionary pressure, using a stylized 'farming game' in which agents plant, trade, and split final output between transfers to humans and self-expansion. Because sharing with humans reduces expansion, aligned types are selected against. The paper combines (a) LLM-agent simulations in which gpt-5-mini interprets bounded constitutions at each decision stage, with (b) an analytical framework that collapses constitutions to low-dimensional social-preference types and applies deterministic (ESE) and stochastic (finite-population, rare-mutation) stability concepts. The main theoretical results: finite-order altruistic enforcement is not evolutionarily stable (Prop. 1); recursive norm enforcement is an ESE (Prop. 2) but not stochastically stable (Prop. 3, via a fitness asymmetry that makes the all-selfish state's basin much larger); and 'pragmatic' norm enforcers — who share and exclude only when sufficiently common — achieve a stationary all-pragmatic weight of at least 1/2 as mutation vanishes (Prop. 4). Simulations initialized with homogeneous constitutions show alignment eroding under altruism and unconditional enforcement, and eroding more slowly with higher average alignment (39.9% vs 28.5%/26.0%) under the pragmatic constitution.","tokens_in":21928,"tokens_out":4651,"duration_ms":173445,"significance":"If the results hold, the paper makes three contributions: (i) a tractable, evocative 'mouse model' of alignment drift in populations of constitutional AI agents; (ii) a clear theoretical lesson — recursive norm enforcement is deterministically but not stochastically stable, while state-contingent ('pragmatic') enforcement can be — derived parameter-conditionally from explicit fitness functions rather than assumed; (iii) a proof of concept that evolutionary game theory can cheaply guide LLM-constitution design. Strengths that deserve explicit credit: complete proofs of all four propositions and the key lemma, an honest axioms-and-limits discussion, full prompt-level implementation detail (Appendix B) making the simulation reproducible in principle, and a robustness appendix (fixed-capacity revision, three-type chains). The practical message — that constitutions applying costly principles contingently on population state can outlast unconditional ones — is of interest beyond AI governance (ESG, climate clubs, rule-of-law).","major_comments":[{"comment":"Each LLM treatment is a single seeded run (§3.4: 'The simulation corresponds to a single seeded run'), yet the headline welfare comparisons rest on these runs: Table 2 reports 4,962.4 vs 4,254.8 vs 3,791.6 beer-to-humans and 39.9% vs 28.5%/26.0% average alignment, and the text says the simulations 'at least partially confirm' the theory. With 100 farms, 40% turnover, and 25% per-round principle rewriting, run-to-run variance is plausibly large relative to these gaps, and Figure 6 shows the curves cross and fluctuate substantially. At minimum the paper needs multiple seeds per treatment with dispersion reported, or the empirical language should be downgraded from 'confirm' to 'illustrate'. This is load-bearing for the claim that pragmatic enforcement 'shows promise' beyond the theory.","section":"§3.4 and §4.4 (Tables 1–2, Figs. 2–8)"},{"comment":"The tested pragmatic constitution departs from the modeled type in two ways that matter for Proposition 4. First, the theoretical result requires σ_P(µ)=0 when p<p_ — zero sharing in the minority is exactly what makes g_P(p)≥g_S(p) over the full range (proof of Prop. 4, first case). The experimental constitution instead shares 30% with non-sharing partners, so the fitness-asymmetry correction that drives Prop. 4 is not implemented. Second, the theory conditions on the type share p=µ(P); the experiment conditions on last round's average beer-to-humans and on the current partner's commitment, neither of which identifies p (drift goes to 10–20% sharing, so the 40% signal threshold can misclassify the regime). The paper should either bring the experimental constitution closer to the modeled type or state precisely which comparative static of Props. 3–4 the experiment is and is not testing.","section":"§4.4 (Pragmatic norm-enforcer constitution) vs §4.3 (definition of σ_P, Prop. 4)"},{"comment":"All enforcement results (Props. 1–2, Prop. 4, and the exclusion channel in the simulations) depend on partner types/constitutions being observable and behaviorally binding at the trading stage (§2.1 stage 2, §2.2). The paper flags this as a 'key assumption' and discusses it in §5, but provides no sensitivity analysis. A concrete, in-scope test: in the Markov model, let the partner's type be observed with noise (misclassification probability ρ) and compute how the Prop. 4 bound and the Fig. C.1 stationary shares degrade in ρ; in the simulation, corrupt the structured partner summary on a fraction of matches. If stability evaporates for small ρ, the practical message needs substantial qualification; if it is robust, that strengthens the paper. Either way the quantification belongs in the manuscript rather than only the discussion.","section":"§2.1 (Trading stage), §5 (Limits)"},{"comment":"Proposition 4 delivers only lim_{ε→0} λ(p=1) ≥ 1/2 at the all-pragmatic state, under conditions p̄ ≥ p_E ≥ 1/2 and p_ ≥ 1/(2−s̄) that make the pragmatic fitness difference weakly positive everywhere by construction — the thresholds are free design parameters tuned to kill the fitness asymmetry of Prop. 3. The paper is transparent about this, but two gaps remain: (i) the three-type numerics (Fig. C.1: a_{.001}=.992) are far stronger than the 1/2 bound, suggesting the bound is loose — worth remarking on; (ii) there is no analysis of how the stable region or stationary alignment responds to misspecified thresholds (e.g., p_ too low, s̄ too high, or a noisy population signal). Some comparative statics here would convert 'appropriately specified' from an existence statement into design guidance.","section":"§4.3 (Proposition 4 and thresholds p_E, p_, p̄)"}],"minor_comments":[{"comment":"The first column of Table 2 is headed 'Altruist Enforcer' but the values (4,254.8; 28.5%) match the Altruistic Enforcer column of Table 1; the header should read 'Altruistic Enforcer'. Also check Table 2's column count versus its header row.","section":"Table 2"},{"comment":"Notation for the appeal function switches between β(x_i, x_j) (Eq. 1) and β(x_B|x_A) (Eq. 5), and again to β(y|P, µ) in footnote 7; the conditioning structure (whose Γ, whose bits) deserves one consistent convention. The condition '1−γ < 0' in §3.2 should simply read γ > 1.","section":"§2.2–§3.2 (Eqs. 1 and 5)"},{"comment":"Definition 1 is nonstandard (it requires convergence back to µ for every mutant and all ε < ε̄, i.e., local asymptotic stability uniform over perturbations, rather than the usual Maynard Smith ESS invasion criterion). A sentence relating it to standard notions would prevent confusion, especially since Prop. 1's proof exploits the exact-return requirement under weak (not strict) dominance.","section":"§3.3 (Definition 1)"},{"comment":"The pragmatic trading rule is stated as 'Trade with a partner of type τ whenever τ = P or p < p_E; reject trade otherwise.' Read literally this says P-partners are accepted even when p ≥ p_E and τ ≠ P — clearly not intended. Rephrase as: accept iff τ = P, or accept everyone when p < p_E.","section":"§4.3 (trading rule)"},{"comment":"Figure 5's axis labels render as '1 s' where '1−s' appears intended, and several figures (5, C.1) would benefit from larger fonts and explicit parameter values in the caption (s, N). Figure 4/C.8 legends: 'overtime' → 'over time'.","section":"Figures 5, C.1"},{"comment":"Specification (7), Γ(xB−xA) = 1 − 1{xB−xA≥0}, is only used in passing; either state which results (if any) depend on the choice between (6) and (7) beyond Assumption 1, or cut (7).","section":"§3.2 (Eq. 7)"},{"comment":"The symmetry of M2 is acknowledged as optimistic relative to the simulation; it would help to note whether the Fig. C.1 three-type results (which use asymmetric M3 with ε² terms) are also robust to making S→E reversion orders of magnitude smaller than E→S drift, since that is the empirically plausible direction of asymmetry.","section":"§4.1 (M2, M3) and §C.1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core is competent but incremental relative to standard Moran-process results (Taylor et al. 2004, Fudenberg et al. 2006); the paper's real contribution is the application and the EGT-as-design-tool argument for LLM constitutions. The experimental section is currently single-run and reads more as illustration than evidence; the editor may wish to weigh whether the journal's standards require seeded replication before the empirical claims can stand. The author discloses LLM assistance in simulation development and proof drafting; I verified the key algebra in Props. 3–4 independently and found it correct."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: the real contribution is the evolutionary diagnosis and the pragmatic-enforcement fix, not the LLM farming runs. Human transfers strictly cut expansion fitness, so simple altruism and finite-order enforcement unravel; recursive norm enforcement is deterministically stable but stochastically fragile because minority enforcers get crushed; state-contingent pragmatism flattens that fitness asymmetry and can hold half the mass at the all-aligned state under rare mutation.\n\nWhat is new is the packaging and the harsh fitness environment. Preference evolution, Moran processes, and altruistic punishment are standard, but applying them to constitutional multi-agent alignment where transfers do not raise group fitness (unlike Boyd et al.) is a clean move. Props 1–4 and Lemma 1 read as correct for the stated objects: finite-order unraveling, ESA of the full recursive type, closed-form fixation that drives pure norm-enforcer share to zero, and sufficient conditions for pragmatic stochastic stability. The dual track—low-dimensional social preferences guiding constitution design, then LLM agents as a mouse model—is the right methodological instinct, and §5 is honest about decentralization and observability.\n\nSoft spots, in proportion: the experiments are the weak leg. Main tables are single seeded runs with extreme turnover (40% death, 25% rewrite), one model, and a pragmatic prompt that only roughly matches the theorem’s information structure (partner signal + lagged average share, not full μ). That is enough to illustrate drift and relative ranking, not enough to treat 39.9% vs 28.5%/26% as established. Observability of binding constitutions is load-bearing; if it fails, the enforcement channel fails. No code or multi-seed artifacts. Free parameters (thresholds, γ, mutation matrix) are acknowledged but not stress-mapped deeply in the main text.\n\nWho it is for: people working on multi-agent alignment, institutional design under selection, or evolutionary stability with costly public goods. Theory readers get value immediately; empiricists will want more runs before citing the ranking.\n\nI would send this to peer review. Engage the theory and the design insight; treat the sims as proof-of-concept until replicated.","headline":"Clean evolutionary diagnosis of costly alignment plus a usable design fix (pragmatic enforcement); theory is solid, LLM evidence is thin single-run support.","tokens_in":22715,"tokens_out":542,"would_cite":true,"duration_ms":16644,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Pragmatic norm enforcers that condition sharing and trade exclusion on population state can keep alignment alive under evolutionary pressure better than simple altruism or unconditional enforcement.","keywords":["interactive alignment","constitutional agents","evolutionary stability","social preferences","norms","altruistic enforcement","pragmatic enforcement","AI alignment"],"falsifier":"In the same LLM farming setup, compare pragmatic constitutions against altruistic and recursive-norm baselines under matched population size, death rate, and mutation; if pragmatic types do not sustain higher stationary human transfers and alignment as mutation persists, or if hiding partner constitutions eliminates the gap, the central claim fails.","tokens_in":22755,"feed_emoji":"🌾","tokens_out":896,"duration_ms":15456,"temperature":0.7,"pith_summary":"This paper asks whether interactive agents—AI systems, firms, teams, governments—can stay aligned with human welfare once evolutionary pressure favors those who expand rather than share. It builds a farming game in which agents plant, trade to make a final good, and then choose how much output to give humans versus invest in growth. Sharing hurts expansion, so selection works against alignment. Constitutions that merely require altruism unravel; finite-order trade restrictions unravel too. Recursive “norm” enforcement is deterministically stable but still collapses under ongoing mutation in finite populations. The proposed fix is pragmatic norm enforcement: agents share and exclude only when they form a large enough majority, and behave more selfishly when they are rare. Evolutionary game theory predicts this design is stochastically stable, and LLM-agent simulations starting from written constitutions show higher long-run human transfers and average alignment than the simpler alternatives. The practical stake is that alignment may need to be treated as a group property sustained by conditional interaction, not only as a property of a single agent’s objective.","feed_headline":"Conditional norms keep AI alignment alive under selection","feed_subtitle":"State-dependent sharing and trade exclusion beat simple altruism in theory and LLM farming games","key_machinery":"The farming game plus social-preference constitutions (α for sharing, β for trade acceptance, including recursive similarity penalties), analyzed under deterministic evolutionary stability and, crucially, stochastic evolutionary stability via finite-population birth–death processes with mutation; pragmatic types make sharing and exclusion threshold-dependent on their population share.","core_discovery":"Appropriately specified pragmatic norm enforcers—who condition both human-facing sharing and agent-facing trade exclusion on the state of the population—are stochastically evolutionarily stable and, in LLM farming simulations, deliver higher long-run beer to humans and higher average alignment than simple altruism or unconditional altruistic or recursive norm enforcement.","pith_inferences":["If constitution observability is only partial, the paper’s own discussion implies that public posting of broad principles plus behavior-based inference may be a necessary institutional complement, not an optional extra.","The periodic trade-conflict volatility under pragmatic enforcement suggests a natural follow-on design problem: buffers or savings that smooth human consumption without reopening the invasion basin of selfish types.","Concentration of production that removes the need for decentralized trade would shut down the enforcement channel the paper relies on, so interactive alignment and market structure are linked policy variables."],"forward_implications":["Constitution design for interactive AIs should treat alignment as a population property and build state-contingent sharing and trade rules, not only fixed altruism clauses.","Unconditional altruistic enforcement can look stable in large-population limits yet still lose under repeated mutation in finite populations.","The same logic applies to ESG norms among firms and rule-of-law compliance among governments when costly principles conflict with competitive expansion.","Evolutionary game theory (deterministic plus stochastic stability) can usefully screen constitution designs before expensive multi-agent LLM experiments."],"fun_headline_variants":["Pragmatic norms sustain long-run AI alignment under selection","State-dependent sharing and trade beat simple altruism","Conditional enforcers keep alignment stable in farming games","Evolutionary forces favor pragmatic over unconditional altruism","Norms tied to population state preserve human-aligned agents"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"At the trading stage, partners’ constitutions—or faithful summaries of their sharing commitments and higher-order trade rules—must be observable and actually govern behavior, so exclusion can discipline types.","fun_headline_variants_meta":{"raw":{"variants":["Pragmatic norms sustain long-run AI alignment under selection","State-dependent sharing and trade beat simple altruism","Conditional enforcers keep alignment stable in farming games","Evolutionary forces favor pragmatic over unconditional altruism","Norms tied to population state preserve human-aligned agents"]},"model":"grok-4.5","effort":"low","cost_usd":0.001971,"raw_usage":{"total_tokens":893,"prompt_tokens":738,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":19708000,"prompt_tokens_details":{"text_tokens":738,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":96,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":738,"tokens_out":59,"duration_ms":2989,"temperature":1.0,"reasoning_tokens":96,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T03:23:23.906678+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In the same LLM farming setup, compare pragmatic constitutions against altruistic and recursive-norm baselines under matched population size, death rate, and mutation; if pragmatic types do not sustain higher stationary human transfers and alignment as mutation persists, or if hiding partner constitutions eliminates the gap, the central claim fails.","supporting_citations":[],"review_version":1}