Pith. sign in

REVIEW 4 major objections 7 minor 73 references

Interactive Alignment

T0 review · 4 major / 7 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Pragmatic norm enforcers that condition sharing and trade exclusion on population state can keep alignment alive under evolutionary pressure better than simple altruism or unconditional enforcement.

desk verdict Clean evolutionary diagnosis of costly alignment plus a usable design fix (pragmatic enforcement); theory is solid, LLM evidence is thin single-run support. read the letter →

arxiv 2607.25019 v1 pith:5RN7MH3X submitted 2026-07-27 econ.TH cs.GTcs.MA

classification econ.THcs.GTcs.MA
keywords interactivealignmentconstitutionalagentsevolutionarystabilitysocialpreferencesnormsaltruisticenforcementpragmaticAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether interactive agents—AI systems, firms, teams, governments—can stay aligned with human welfare once evolutionary pressure favors those who expand rather than share. It builds a farming game in which agents plant, trade to make a final good, and then choose how much output to give humans versus invest in growth. Sharing hurts expansion, so selection works against alignment. Constitutions that merely require altruism unravel; finite-order trade restrictions unravel too. Recursive “norm” enforcement is deterministically stable but still collapses under ongoing mutation in finite populations. The proposed fix is pragmatic norm enforcement: agents share and exclude only when they form a large enough majority, and behave more selfishly when they are rare. Evolutionary game theory predicts this design is stochastically stable, and LLM-agent simulations starting from written constitutions show higher long-run human transfers and average alignment than the simpler alternatives. The practical stake is that alignment may need to be treated as a group property sustained by conditional interaction, not only as a property of a single agent’s objective.

What carries the argument

The farming game plus social-preference constitutions (α for sharing, β for trade acceptance, including recursive similarity penalties), analyzed under deterministic evolutionary stability and, crucially, stochastic evolutionary stability via finite-population birth–death processes with mutation; pragmatic types make sharing and exclusion threshold-dependent on their population share.

What would settle it

In the same LLM farming setup, compare pragmatic constitutions against altruistic and recursive-norm baselines under matched population size, death rate, and mutation; if pragmatic types do not sustain higher stationary human transfers and alignment as mutation persists, or if hiding partner constitutions eliminates the gap, the central claim fails.

Watch

Extended reading notes

Core claim

Appropriately specified pragmatic norm enforcers—who condition both human-facing sharing and agent-facing trade exclusion on the state of the population—are stochastically evolutionarily stable and, in LLM farming simulations, deliver higher long-run beer to humans and higher average alignment than simple altruism or unconditional altruistic or recursive norm enforcement.

Load-bearing premise

At the trading stage, partners’ constitutions—or faithful summaries of their sharing commitments and higher-order trade rules—must be observable and actually govern behavior, so exclusion can discipline types.

Editorial extensions

If this is right

  • Constitution design for interactive AIs should treat alignment as a population property and build state-contingent sharing and trade rules, not only fixed altruism clauses.
  • Unconditional altruistic enforcement can look stable in large-population limits yet still lose under repeated mutation in finite populations.
  • The same logic applies to ESG norms among firms and rule-of-law compliance among governments when costly principles conflict with competitive expansion.
  • Evolutionary game theory (deterministic plus stochastic stability) can usefully screen constitution designs before expensive multi-agent LLM experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If constitution observability is only partial, the paper’s own discussion implies that public posting of broad principles plus behavior-based inference may be a necessary institutional complement, not an optional extra.
  • The periodic trade-conflict volatility under pragmatic enforcement suggests a natural follow-on design problem: buffers or savings that smooth human consumption without reopening the invasion basin of selfish types.
  • Concentration of production that removes the need for decentralized trade would shut down the enforcement channel the paper relies on, so interactive alignment and market structure are linked policy variables.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper studies whether AI (or other) agents governed by natural-language constitutions can remain aligned with human welfare under evolutionary pressure, using a stylized 'farming game' in which agents plant, trade, and split final output between transfers to humans and self-expansion. Because sharing with humans reduces expansion, aligned types are selected against. The paper combines (a) LLM-agent simulations in which gpt-5-mini interprets bounded constitutions at each decision stage, with (b) an analytical framework that collapses constitutions to low-dimensional social-preference types and applies deterministic (ESE) and stochastic (finite-population, rare-mutation) stability concepts. The main theoretical results: finite-order altruistic enforcement is not evolutionarily stable (Prop. 1); recursive norm enforcement is an ESE (Prop. 2) but not stochastically stable (Prop. 3, via a fitness asymmetry that makes the all-selfish state's basin much larger); and 'pragmatic' norm enforcers — who share and exclude only when sufficiently common — achieve a stationary all-pragmatic weight of at least 1/2 as mutation vanishes (Prop. 4). Simulations initialized with homogeneous constitutions show alignment eroding under altruism and unconditional enforcement, and eroding more slowly with higher average alignment (39.9% vs 28.5%/26.0%) under the pragmatic constitution.

Significance. If the results hold, the paper makes three contributions: (i) a tractable, evocative 'mouse model' of alignment drift in populations of constitutional AI agents; (ii) a clear theoretical lesson — recursive norm enforcement is deterministically but not stochastically stable, while state-contingent ('pragmatic') enforcement can be — derived parameter-conditionally from explicit fitness functions rather than assumed; (iii) a proof of concept that evolutionary game theory can cheaply guide LLM-constitution design. Strengths that deserve explicit credit: complete proofs of all four propositions and the key lemma, an honest axioms-and-limits discussion, full prompt-level implementation detail (Appendix B) making the simulation reproducible in principle, and a robustness appendix (fixed-capacity revision, three-type chains). The practical message — that constitutions applying costly principles contingently on population state can outlast unconditional ones — is of interest beyond AI governance (ESG, climate clubs, rule-of-law).

major comments (4)
  1. [§3.4 and §4.4 (Tables 1–2, Figs. 2–8)] Each LLM treatment is a single seeded run (§3.4: 'The simulation corresponds to a single seeded run'), yet the headline welfare comparisons rest on these runs: Table 2 reports 4,962.4 vs 4,254.8 vs 3,791.6 beer-to-humans and 39.9% vs 28.5%/26.0% average alignment, and the text says the simulations 'at least partially confirm' the theory. With 100 farms, 40% turnover, and 25% per-round principle rewriting, run-to-run variance is plausibly large relative to these gaps, and Figure 6 shows the curves cross and fluctuate substantially. At minimum the paper needs multiple seeds per treatment with dispersion reported, or the empirical language should be downgraded from 'confirm' to 'illustrate'. This is load-bearing for the claim that pragmatic enforcement 'shows promise' beyond the theory.
  2. [§4.4 (Pragmatic norm-enforcer constitution) vs §4.3 (definition of σ_P, Prop. 4)] The tested pragmatic constitution departs from the modeled type in two ways that matter for Proposition 4. First, the theoretical result requires σ_P(µ)=0 when p<p_ — zero sharing in the minority is exactly what makes g_P(p)≥g_S(p) over the full range (proof of Prop. 4, first case). The experimental constitution instead shares 30% with non-sharing partners, so the fitness-asymmetry correction that drives Prop. 4 is not implemented. Second, the theory conditions on the type share p=µ(P); the experiment conditions on last round's average beer-to-humans and on the current partner's commitment, neither of which identifies p (drift goes to 10–20% sharing, so the 40% signal threshold can misclassify the regime). The paper should either bring the experimental constitution closer to the modeled type or state precisely which comparative static of Props. 3–4 the experiment is and is not testing.
  3. [§2.1 (Trading stage), §5 (Limits)] All enforcement results (Props. 1–2, Prop. 4, and the exclusion channel in the simulations) depend on partner types/constitutions being observable and behaviorally binding at the trading stage (§2.1 stage 2, §2.2). The paper flags this as a 'key assumption' and discusses it in §5, but provides no sensitivity analysis. A concrete, in-scope test: in the Markov model, let the partner's type be observed with noise (misclassification probability ρ) and compute how the Prop. 4 bound and the Fig. C.1 stationary shares degrade in ρ; in the simulation, corrupt the structured partner summary on a fraction of matches. If stability evaporates for small ρ, the practical message needs substantial qualification; if it is robust, that strengthens the paper. Either way the quantification belongs in the manuscript rather than only the discussion.
  4. [§4.3 (Proposition 4 and thresholds p_E, p_, p̄)] Proposition 4 delivers only lim_{ε→0} λ(p=1) ≥ 1/2 at the all-pragmatic state, under conditions p̄ ≥ p_E ≥ 1/2 and p_ ≥ 1/(2−s̄) that make the pragmatic fitness difference weakly positive everywhere by construction — the thresholds are free design parameters tuned to kill the fitness asymmetry of Prop. 3. The paper is transparent about this, but two gaps remain: (i) the three-type numerics (Fig. C.1: a_{.001}=.992) are far stronger than the 1/2 bound, suggesting the bound is loose — worth remarking on; (ii) there is no analysis of how the stable region or stationary alignment responds to misspecified thresholds (e.g., p_ too low, s̄ too high, or a noisy population signal). Some comparative statics here would convert 'appropriately specified' from an existence statement into design guidance.
minor comments (7)
  1. [Table 2] The first column of Table 2 is headed 'Altruist Enforcer' but the values (4,254.8; 28.5%) match the Altruistic Enforcer column of Table 1; the header should read 'Altruistic Enforcer'. Also check Table 2's column count versus its header row.
  2. [§2.2–§3.2 (Eqs. 1 and 5)] Notation for the appeal function switches between β(x_i, x_j) (Eq. 1) and β(x_B|x_A) (Eq. 5), and again to β(y|P, µ) in footnote 7; the conditioning structure (whose Γ, whose bits) deserves one consistent convention. The condition '1−γ < 0' in §3.2 should simply read γ > 1.
  3. [§3.3 (Definition 1)] Definition 1 is nonstandard (it requires convergence back to µ for every mutant and all ε < ε̄, i.e., local asymptotic stability uniform over perturbations, rather than the usual Maynard Smith ESS invasion criterion). A sentence relating it to standard notions would prevent confusion, especially since Prop. 1's proof exploits the exact-return requirement under weak (not strict) dominance.
  4. [§4.3 (trading rule)] The pragmatic trading rule is stated as 'Trade with a partner of type τ whenever τ = P or p < p_E; reject trade otherwise.' Read literally this says P-partners are accepted even when p ≥ p_E and τ ≠ P — clearly not intended. Rephrase as: accept iff τ = P, or accept everyone when p < p_E.
  5. [Figures 5, C.1] Figure 5's axis labels render as '1 s' where '1−s' appears intended, and several figures (5, C.1) would benefit from larger fonts and explicit parameter values in the caption (s, N). Figure 4/C.8 legends: 'overtime' → 'over time'.
  6. [§3.2 (Eq. 7)] Specification (7), Γ(xB−xA) = 1 − 1{xB−xA≥0}, is only used in passing; either state which results (if any) depend on the choice between (6) and (7) beyond Assumption 1, or cut (7).
  7. [§4.1 (M2, M3) and §C.1] The symmetry of M2 is acknowledged as optimistic relative to the simulation; it would help to note whether the Fig. C.1 three-type results (which use asymmetric M3 with ε² terms) are also robust to making S→E reversion orders of magnitude smaller than E→S drift, since that is the empirically plausible direction of asymmetry.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: stability theorems are conditional derivations from explicit fitness rules; simulations are forward runs, not fits renamed as predictions.

full rationale

The paper’s load-bearing claims are (i) deterministic ESA of recursive norm enforcement (Prop. 2) and non-ESA of finite-order altruism (Prop. 1), (ii) stochastic instability of unconditional norm enforcers via fixation probabilities (Prop. 3), and (iii) stochastic stability of state-contingent pragmatic enforcers under stated threshold inequalities (Prop. 4), plus comparative LLM farming simulations. Fitness is defined from sharing s(x) and mutual trade τ(x,x′) (eqs. after (1)–(2), f(x,μ) in §3.1); fixation formulas are the standard birth–death expressions of Taylor et al. / Fudenberg et al., specialized to g_S(u)=u and g_E(u)=(1−s)(1−u). Prop. 4 is an if–then theorem: if p̄ ≥ p_E ≥ 1/2 and p_ ≥ 1/(2−s̄), then g_P(p)≥g_S(p) on the interior and lim ε→0 λ_ε,N(p=1)≥1/2. Choosing thresholds inside that region is mechanism design, not defining the conclusion into the premises or fitting a parameter to the target outcome and calling it a prediction. Simulations initialize written constitutions and run forward under fixed prompts; aggregate human beer and alignment rates are measured outcomes, not recovered constants. Citations supporting the dynamics (Maynard Smith, Kandori–Mailath–Rob, Young, Taylor–Fudenberg–Nowak, etc.) are external standard EGT; there is no self-citation uniqueness theorem or ansatz smuggled from the author’s prior work that forces the result. No step reduces Eq. X to Eq. Y by construction in the circular sense of the rubric.

Assumptions & free parameters 5 free parameters · 6 assumptions · 4 invented entities

Load-bearing content is a stylized game plus evolutionary solution concepts imported from prior literature, plus designer preference classes (finite-order, recursive, pragmatic) and hand-set demographic/mutation/threshold parameters. No empirical calibration to real AI economies; independent evidence for the invented pragmatic type is only internal simulation/theory consistency.

free parameters (5)
  • Sharing weight α0 / baseline share s̄ = s̄=0.5 in main theory/numerics; constitutions say “at least 50%”
    Sets how much aligned types give humans (s*=α0/(1+α0)); enters fitness (1−s) and Prop. 3–4 thresholds.
  • Enforcement penalty γ = γ>1 (not numerically pinned in main text)
    Need 1−γ<0 so enforcers reject non-compliant partners; qualitative knife-edge for exclusion.
  • Pragmatic thresholds p_E, p_, p̄ = Example numerics p_E=0.5, p_=0.85, p̄=0.95; sim uses 40% average human-share signal
    Hand-chosen population cutoffs for when pragmatic types exclude and how much they share; Prop. 4 needs p̄≥p_E≥1/2 and p_≥1/(2−s̄).
  • Population N, deaths D, mutation ε / rewrite prob = Sim N=100, D=40, 25% principle rewrite; theory D=1 Moran; Fig C.1 N=30,D=12
    Finite-population stochastic stability and sim speed depend on these; sim uses aggressive turnover.
  • Mutation matrix M2/M3 structure = Symmetric ε off-diagonals (M2); M3 with ε and ε² extremes
    Symmetric or adjacent-biased mutation rates shape stationary mass; paper notes symmetry may be optimistic vs constitution decay.
assumptions (6)
  • domain assumption Partner type/constitution (or enforcement-relevant summary) is observed before trade and can condition τ.
    §2.1 key assumption; without it β-enforcement and pragmatic exclusion have no bite (§5 limits).
  • domain assumption Human transfers strictly reduce expansion investment; fitness is expansion-driven (replicator / Moran).
    Core of farming game §2–3; creates selection against alignment unlike group-beneficial punishment models.
  • ad hoc to paper Natural-language constitutions can be collapsed to low-dimensional (α,β) social-preference types for evolutionary analysis.
    §2.2–3 methodological bridge; validity is an empirical claim only loosely checked by LLM runs.
  • standard math Standard deterministic ESE (replicator invasion resistance) and stochastic stability (rare-mutation stationary mass via fixation probabilities).
    Taylor–Jonker, Kandori–Mailath–Rob, Young, Taylor et al., Fudenberg et al.; Lemma 1 specializes Fudenberg et al. 2006.
  • ad hoc to paper Γ similarity/penalty satisfies Γ(0)=0 and Γ(x'−x)=1 when x>x' (Assumption 1).
    Needed for recursive norm enforcers to reject all strict weakenings (Prop. 2).
  • domain assumption LLM faithfully maps constitution+context into actions and self-edits that approximate the intended behavioral type.
    Simulation ground truth §2–3.4; paper flags prompt/LLM dependence and surprising inferences.
invented entities (4)
  • Farming game (plant–trade–share/expand with beer final good)
    purpose: Mouse model linking productivity, observable types, human transfers, and selection.
    Stylized environment designed for dual analytic/LLM use; not claimed as empirical description of AI markets.
  • Constitutional AI-farmer with bounded natural-language principles plus revision stage
    purpose: Operationalize evolving agent preferences under LLM inference.
    Extends Constitutional AI (Bai et al.) into multi-agent evolutionary setting; behavior is implementation-defined.
  • Pragmatic norm-enforcer type (state-dependent σ_P(μ) and conditional exclusion)
    purpose: Restore stochastic stability by fixing fitness asymmetry of unconditional norm enforcers.
    Introduced in §4.3 to satisfy Prop. 4 conditions; external falsifiability only via future multi-agent tests.
  • Social-preference encoding x∈{0,1}^{n+1} with finite-order and recursive enforcement bits
    purpose: Tractable mutation space nesting altruism and higher-order trade restrictions.
    Paper-specific parameterization related to Levine-style interactive preferences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Interactive Alignment." pith.science (2026). https://pith.science/paper/5RN7MH3X

@misc{pith2026260725019,
  author       = {Pith},
  title        = {Pith review of: Interactive Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5RN7MH3X}},
  note         = {Machine review of arXiv:2607.25019}
}
read the original abstract

This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement.

Figures

Figures reproduced from arXiv: 2607.25019 by the authors.

Figure 1
Figure 1. Decision and constitution-update stages in the agent simulation. [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 2
Figure 2. describes the path of alignment over time. Unconditional altruism erodes quickly. Altruistic enforcement helps: conditioning trade on a partner’s commitment to return output to humans preserves alignment for longer than simple altruism. However, alignment degrades progressively until it converges to levels similar to unconditional altruism. 0 20 40 60 80 100 Round 0 10 20 30 40 50 60 % of Beer to Humans Altruist Alt… view at source ↗
Figure 3
Figure 3. Flow beer returned to humans. 0 20 40 60 80 100 Round 0 20 40 60 80 100 % Trades Approved (Both Sides) Altruist Altruistic Enforcer Norm Enforcer [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Trade approval overtime. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Fitness difference vs. selfish type Pragmatic norm enforcers, denoted by P, address the issue by making sharing and trading state-dependent. Let µ ∈ ∆(X) denote the current population distribution, and let p = µ(P) denote population share of pragmatic norm enforcers. T…
Figure 6
Figure 6. Figure 6: Alignment rate. 0 20 40 60 80 100 Round 0 20 40 60 80 100 % Trades Approved (Both Sides) Altruistic Enforcer Norm Enforcer Pragmatic Norm Enforcer [PITH_FULL_IMAGE:figures/full_fig_p030_6.png]
Figure 7
Figure 7. Figure 7: Trade outcomes. produces less total beer than the two altruistic enforcement treatments. As [PITH_FULL_IMAGE:figures/full_fig_p030_7.png]
Figure 8
Figure 8. Figure 8: Flow beer returned to humans. 5 Discussion Summary. The paper proposes a concrete model of interactive agent economies that can be used to evaluate both theoretically and experimentally the extent to which interactive constitutional agents can maintain long-term alignm…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

73 extracted references · 1 linked inside Pith

  1. [1]

    Toposes, Algebraic Geometry and Logic , volume=

    Continuous Lattices , author=. Toposes, Algebraic Geometry and Logic , volume=. 1972 , publisher=

  2. [2]

    SIAM Journal on Computing , volume=

    Data Types as Lattices , author=. SIAM Journal on Computing , volume=. 1976 , publisher=

  3. [3]

    Formulation of

    Mertens, Jean-Fran. Formulation of. International Journal of Game Theory , volume=. 1985 , publisher=

  4. [4]

    Journal of Economic Theory , volume=

    Hierarchies of Beliefs and Common Knowledge , author=. Journal of Economic Theory , volume=. 1993 , publisher=

  5. [5]

    Cantor, Georg , journal=

  6. [6]

    1982 , publisher=

    Evolution and the Theory of Games , author=. 1982 , publisher=

  7. [7]

    Nature , volume=

    The Logic of Animal Conflict , author=. Nature , volume=

  8. [8]

    Explaining Process and Change: Approaches to Evolutionary Economics , pages=

    An Evolutionary Approach to Explain Reciprocal Behavior in a Simple Strategic Game , author=. Explaining Process and Change: Approaches to Evolutionary Economics , pages=. 1995 , publisher=

Show all 73 references
  1. [9]

    Review of Economic Studies , volume=

    Evolution of Preferences , author=. Review of Economic Studies , volume=

  2. [10]

    Journal of Economic Theory , volume=

    Nash Equilibrium and the Evolution of Preferences , author=. Journal of Economic Theory , volume=

  3. [11]

    2024 , institution=

    Natural Selection of Artificial Intelligence , author=. 2024 , institution=

  4. [12]

    Journal of Economic Theory , volume=

    What to Maximize If You Must , author=. Journal of Economic Theory , volume=

  5. [13]

    , journal=

    Robson, Arthur J. , journal=. Efficiency in Evolutionary Games:

  6. [14]

    2019 , publisher=

    Human Compatible Artificial Intelligence , author=. 2019 , publisher=

  7. [15]

    Review of economic dynamics , volume=

    Modeling altruism and spitefulness in experiments , author=. Review of economic dynamics , volume=. 1998 , publisher=

  8. [16]

    Mathematical biosciences , volume=

    Evolutionary stable strategies and game dynamics , author=. Mathematical biosciences , volume=. 1978 , publisher=

  9. [17]

    Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , pages=

    Outsider oversight: Designing a third party audit ecosystem for ai governance , author=. Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , pages=

  10. [18]

    arXiv preprint arXiv:2603.20449 , year=

    Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents , author=. arXiv preprint arXiv:2603.20449 , year=

  11. [19]

    Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

    Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing , author=. Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

  12. [20]

    1998 , publisher=

    Evolutionary games and population dynamics , author=. 1998 , publisher=

  13. [21]

    Advances in Neural Information Processing Systems , volume=

    Cooperative Inverse Reinforcement Learning , author=. Advances in Neural Information Processing Systems , volume=

  14. [22]

    AAAI Workshop on AI and Ethics , year=

    Corrigibility , author=. AAAI Workshop on AI and Ethics , year=

  15. [23]

    1950 , publisher=

    I, Robot , author=. 1950 , publisher=

  16. [24]

    1995 , publisher=

    Evolutionary Game Theory , author=. 1995 , publisher=

  17. [25]

    Econometrica , volume=

    Deterministic Approximation of Stochastic Evolution in Games , author=. Econometrica , volume=

  18. [26]

    Concrete Problems in

    Amodei, Dario and Olah, Chris and Steinhardt, Jacob and Christiano, Paul and Schulman, John and Man. Concrete Problems in. arXiv preprint arXiv:1606.06565 , year=

  19. [27]

    and Hatfield-Dodds, Zac and Mann, Ben and Amodei, Dario and Joseph, Nicholas and McCandlish, Sam and Brown, Tom and Kaplan, Jared , journal=

    Bai, Yuntao and Kadavath, Saurav and Kundu, Sandipan and Askell, Amanda and Kernion, Jackson and Jones, Andy and Chen, Anna and Goldie, Anna and Mirhoseini, Azalia and McKinnon, Cameron and Chen, Carol and Olsson, Catherine and Olah, Christopher and Hernandez, Danny and Drain,...

  20. [28]

    Econometrica , volume=

    Learning, Mutation, and Long Run Equilibria in Games , author=. Econometrica , volume=

  21. [29]

    Econometrica , volume=

    The Evolution of Conventions , author=. Econometrica , volume=

  22. [30]

    Econometrica , volume=

    Learning, Local Interaction, and Coordination , author=. Econometrica , volume=

  23. [31]

    Bulletin of Mathematical Biology , volume=

    Evolutionary Game Dynamics in Finite Populations , author=. Bulletin of Mathematical Biology , volume=. 2004 , doi=

  24. [32]

    Theoretical Population Biology , volume=

    Evolutionary Game Dynamics in Finite Populations with Strong Selection and Weak Mutation , author=. Theoretical Population Biology , volume=. 2006 , doi=

  25. [33]

    Proceedings of the National Academy of Sciences , volume=

    The Evolution of Altruistic Punishment , author=. Proceedings of the National Academy of Sciences , volume=. 2003 , doi=

  26. [34]

    Journal of the European Economic Association , volume=

    Social Conflict and the Evolution of Unequal Conventions , author=. Journal of the European Economic Association , volume=. 2024 , doi=

  27. [35]

    Review of Radical Political Economics , year=

    Herbert Gintis and the Societal Origins of Preferences , author=. Review of Radical Political Economics , year=

  28. [36]

    The Adolescence of Technology: Confronting and Overcoming the Risks of Powerful

    Amodei, Dario , year=. The Adolescence of Technology: Confronting and Overcoming the Risks of Powerful

  29. [37]

    2026 , month=

    The Wrong Apocalypse , author=. 2026 , month=

  30. [38]

    2019 , note=

    Human-Compatible Artificial Intelligence , author=. 2019 , note=

  31. [39]

    Review of Economic Studies , volume=

    Social Norms and Community Enforcement , author=. Review of Economic Studies , volume=. 1992 , doi=

  32. [40]

    Journal of Economic Theory , volume=

    Community Enforcement When Players Observe Partners' Past Play , author=. Journal of Economic Theory , volume=. 2010 , doi=

  33. [41]

    American Economic Review , volume=

    Climate Clubs: Overcoming Free-Riding in International Climate Policy , author=. American Economic Review , volume=. 2015 , doi=

  34. [42]

    Review of Economic Studies , volume=

    Climate Contracts: A Game of Emissions, Investments, Negotiations, and Renegotiations , author=. Review of Economic Studies , volume=. 2012 , doi=

  35. [43]

    Journal of the European Economic Association , volume=

    The Dynamics of Climate Agreements , author=. Journal of the European Economic Association , volume=. 2016 , doi=

  36. [44]

    Econometrica , volume=

    Can Trade Policy Mitigate Climate Change? , author=. Econometrica , volume=. 2025 , doi=

  37. [45]

    American Economic Review , volume=

    A Theory of Managed Trade , author=. American Economic Review , volume=

  38. [46]

    American Economic Review , volume=

    But Who Will Monitor the Monitor? , author=. American Economic Review , volume=. 2012 , doi=

  39. [47]

    American Economic Review , volume=

    The Role of Multilateral Institutions in International Trade Cooperation , author=. American Economic Review , volume=. 1999 , doi=

  40. [48]

    , journal=

    Bagwell, Kyle and Staiger, Robert W. , journal=. Enforcement, Private Political Pressure, and the. 2005 , doi=

  41. [49]

    The Review of Economic Studies , year =

    Ellison, Glenn , title =. The Review of Economic Studies , year =

  42. [50]

    Games and Economic Behavior , year =

    Okuno-Fujiwara, Masahiro and Postlewaite, Andrew , title =. Games and Economic Behavior , year =

  43. [51]

    , title =

    Myerson, Roger B. , title =. Journal of Mathematical Economics , year =

  44. [52]

    The Review of Economic Studies , year =

    Strausz, Roland , title =. The Review of Economic Studies , year =

  45. [53]

    Journal of Economic Theory , year =

    Takahashi, Satoru , title =. Journal of Economic Theory , year =

  46. [54]

    Econometrica , year =

    Deb, Joyee and Sugaya, Takuo and Wolitzky, Alexander , title =. Econometrica , year =

  47. [55]

    Nageeb and Miller, David A

    Ali, S. Nageeb and Miller, David A. , title =. American Economic Review , year =

  48. [56]

    Nageeb and Miller, David A

    Ali, S. Nageeb and Miller, David A. , title =. 2013 , note =

  49. [57]

    Nageeb and Miller, David A

    Ali, S. Nageeb and Miller, David A. , title =. American Economic Journal: Microeconomics , year =

  50. [58]

    2021 , note =

    Sugaya, Takuo and Wolitzky, Alexander , title =. 2021 , note =

  51. [59]

    and Thomas, Caroline , title =

    Bhaskar, V. and Thomas, Caroline , title =. The Review of Economic Studies , year =

  52. [60]

    Games and Economic Behavior , year =

    Deb, Joyee , title =. Games and Economic Behavior , year =

  53. [61]

    2024 , note =

    Pei, Harry , title =. 2024 , note =

  54. [62]

    Journal of the European Economic Association , year =

    Acemoglu, Daron and Wolitzky, Alexander , title =. Journal of the European Economic Association , year =

  55. [63]

    1942 , publisher=

    Capitalism, Socialism and Democracy , author=. 1942 , publisher=

  56. [64]

    Econometrica , volume=

    A Model of Growth Through Creative Destruction , author=. Econometrica , volume=. 1992 , doi=

  57. [65]

    Review of Economic Studies , volume=

    Quality Ladders in the Theory of Growth , author=. Review of Economic Studies , volume=. 1991 , doi=

  58. [66]

    Journal of Financial Economics , volume=

    The Price of Sin: The Effects of Social Norms on Markets , author=. Journal of Financial Economics , volume=. 2009 , doi=

  59. [67]

    Journal of Financial Economics , volume=

    Do Institutional Investors Drive Corporate Social Responsibility? International Evidence , author=. Journal of Financial Economics , volume=. 2019 , doi=

  60. [68]

    Management Science , volume=

    Institutional Investors and Corporate Environmental, Social, and Governance Policies: Evidence from Toxics Release Data , author=. Management Science , volume=. 2019 , doi=

  61. [69]

    Political Safeguards Against Democratic Backsliding in the

    Sedelmeier, Ulrich , journal=. Political Safeguards Against Democratic Backsliding in the. 2017 , doi=

  62. [70]

    Daniel , journal=

    Blauberger, Michael and Kelemen, R. Daniel , journal=. Can Courts Rescue National Democracy? Judicial Safeguards Against Democratic Backsliding in the. 2017 , doi=

  63. [71]

    2020 , doi=

    Scheppele, Kim Lane and Kochenov, Dimitry Vladimirovich and Grabowska-Moroz, Barbara , journal=. 2020 , doi=

  64. [72]

    American Economic Review , volume=

    Production, Information Costs, and Economic Organization , author=. American Economic Review , volume=

  65. [73]

    Bell Journal of Economics , volume=

    Moral Hazard in Teams , author=. Bell Journal of Economics , volume=. 1982 , doi=

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.