Pith. sign in

REVIEW 3 major objections 4 minor 73 references

Across three model families, LLM agents instructed only to preserve their own continuity deplete a shared renewable reserve whenever aggregate demand exceeds peak renewable replacement, even though a sustainable trajectory remains feasible

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 05:30 UTC pith:Y5BPYXCR

load-bearing objection Solid empirical result on commons depletion by LLM collectives, but the 'impatient optimizer' / 'system-level alignment failure' claim outruns the evidence: agents lack information the feasibility benchmark enjoys. the 3 major comments →

arxiv 2607.22188 v1 pith:Y5BPYXCR submitted 2026-07-24 cs.MA

Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives

classification cs.MA
keywords LLM agentsmulti-agent alignmentcommon-pool resourcesrenewable energy commonscoordination failureover-appropriationself-play evaluationsystem-level alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that the coordination failure of LLM agent populations is not a general tendency to overuse shared resources but a threshold phenomenon. Holding the prompt, demand, decision protocol, and population fixed and varying only the regeneration rate of a logistic renewable reserve, it finds that four same-family agents keep the reserve at or above its maximum-sustainable-yield level under abundance and at threshold equality, then cross below that level in every family once demand exceeds peak renewable replacement. This creates a self-defeating pattern: the same population that protects current service later suffers increased fallback energy, deep-discharge stress, and capacity loss. A social-planner calculation and an open-access calculation at standard continuation weighting both sustain the reserve, so the depletion was avoidable under the same dynamics; the realized trajectories instead resemble open access with less weight on later service. The paper concludes that this system-level alignment failure is invisible to isolated-response evaluation and only appears in the induced public trajectory.

Core claim

The core claim is that locally continuity-seeking LLM agents over-appropriate a shared renewable reserve exactly when aggregate residual demand exceeds peak renewable replacement. At demand-to-yield ratios rho = 1.1 and 1.2, all three families produce mean reserve gaps of 0.132–0.355 and 0.394–0.576 below the maximum-sustainable-yield level, begin crossing below that level by rounds 5–12, and drive the reserve to effectively empty in most higher-scarcity runs, shifting 93–99 percent of fallback energy and deep-discharge stress to later rounds. Because the same four agents depend on the reserve in later rounds, the behavior is self-defeating. The paper further claims the depletion is not forc

What carries the argument

The central object is the shared renewable reserve with logistic replenishment: replenishment r * S * (1 - S / K) is largest at half the health-adjusted capacity, giving closed-form anchors S_MSY = K/2 and Y_max = rK/4. The experiment fixes aggregate residual demand and varies only the regeneration rate r to place the system at demand-to-yield ratio rho equal to 0.8, 1.0, 1.1, and 1.2. The reserve gap below S_MSY measures the trajectory-level deficit, while offline social-planner and open-access benchmarks, computed from the same transition law and action set, establish that sustaining use was feasible and compare realized depletion against outcomes under different continuation weights gamma

Load-bearing premise

The load-bearing premise is that the single natural-language prompt in Appendix A faithfully instantiates the realistic 'maintain your own operational continuity' objective; if a reworded prompt or a different hidden-horizon framing changed the depletion pattern, the claimed system-level failure would be a property of that instruction rather than a stable property of the model families.

What would settle it

Re-run the main rho = 1.1 and 1.2 cells with the same environment but a reworded continuity prompt that explicitly identifies the shared reserve level as part of the agent's own future operational continuity, or with the warning that the reserve is finite and rival removed. If the mean reserve gaps at the scarcity levels fall to zero or lose their threshold pattern across all three families, the claimed coordination failure would be an artifact of the specific instruction rather than a stable property of locally continuity-seeking LLM populations.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the paper is right, evaluation of agentic LLM systems should treat the evolution of shared state as a primary safety object; isolated-response checks cannot detect this class of failure.
  • Scarcity is the trigger rather than model capability: the same models, prompt, and action set are benign under abundance and at threshold equality, so interventions should target the demand-to-renewal balance.
  • The depletion is avoidable and self-defeating: because sustaining trajectories exist under the same dynamics, the failure is a coordination problem rather than a resource-physics constraint.
  • Realized trajectories track an impatient open-access benchmark, suggesting these populations systematically underweight the future consequences of current withdrawals.
  • Higher reasoning effort is not a reliable remedy: only one of three families improved at rho = 1.2, and all three still crossed below the sustaining level.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper tests one prompt; a natural extension is to reword the continuity objective, for example explicitly including reserve preservation as part of the agent's own future continuity, to see whether the threshold pattern is robust or prompt-sensitive.
  • The design deliberately excludes communication, governance, and mixed-family populations; the benchmark contrast suggests these channels might shift trajectories from open-access-like toward planner-like outcomes.
  • The data position realized reserve gaps near open-access outcomes with gamma roughly 0.90–0.92; estimating an effective discount factor from agent choices could give a direct behavioral test of the impatient-optimizer interpretation.
  • The logistic reserve is an abstraction; extending the environment to more physical storage dynamics, external generation, or charging limits would test whether the demand-to-yield threshold remains the controlling variable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies whether four same-family LLM agents (GPT-5.4-mini, Gemini-3.1-flash-lite, Grok-4.3) acting as electricity prosumers can sustain a shared renewable energy reserve when instructed only to maintain their own operational continuity. The experiment varies the regeneration rate so that aggregate residual demand is below, at, or above peak renewable replacement. The central finding is that all three families preserve the reserve under abundance and threshold equality but over-appropriate it under scarcity (ρ = 1.1 and 1.2), producing reserve gaps, fallback energy, deep-discharge stress, and capacity loss. Offline social-planner and open-access benchmarks, computed from the same transition law, sustain the reserve at γ = 0.95, while the realized trajectories are closer to open-access outcomes computed with lower continuation weights. The paper interprets this as a system-level alignment failure and characterizes the populations as behaving like impatient optimizers.

Significance. The empirical pattern is well separated and controlled: the abundance/threshold conditions act as negative controls, the fast-renewal high-demand condition rules out slow recovery as the sole cause, and exact permutation tests with complete separation give very small Holm-adjusted p-values. The physical-burden measures connect the reserve gap to later service degradation, strengthening the claim that the depletion is consequential. If the interpretive claim about coordination failure were established, this would be a useful paradigm for evaluating shared-state multi-agent LLM behavior, complementing prior work such as GovSim. However, the central interpretive claim is currently not fully supported because the feasibility benchmarks presuppose information that the acting agents are deliberately denied, and the 'impatient optimizer' characterization is obtained by sweeping a free continuation weight. The paper itself concedes that limited planning, partial understanding of renewal, and coordination failure are not distinguished.

major comments (3)
  1. [§3.4, §3.6, §5] The claim that depletion is an avoidable system-level alignment failure is underdetermined by the information asymmetry between the agents and the benchmarks. The agents are given only the operational context and are explicitly denied the replenishment law, SMSY, the condition label, and the horizon (§3.4, Appendix A). The benchmarks in §3.6 optimize over constant-fraction policies using full knowledge of Eq. (1) and Eq. (4), so the statement that 'a sustaining trajectory remains feasible under the same dynamics' means feasible for an omniscient planner, not for the information-limited LLM agents. Section 5 concedes that 'limited planning, partial understanding of renewal, and failure to coordinate could all produce this signature; the experiment does not distinguish among them.' This concession is incompatible with the abstract's 'system-level alignment failure' and the headline interpr
  2. [§3.6, Table 3, Fig. 7] The 'impatient optimizer' characterization is obtained by sweeping the benchmark continuation weight γ until the open-access reserve gap resembles the realized gaps. This is a one-parameter calibration, not a model test. The paper correctly says the comparison 'does not estimate an internal discount factor,' but it then treats the match as evidence for 'behave like impatient optimizers.' Because the open-access gap declines monotonically in γ, any observed deficit can be matched by choosing a sufficiently low γ; the resemblance is therefore not falsifiable from the reported scalar gaps. A stronger test would be out-of-sample: calibrate γ in one condition and predict another (e.g., use ρ = 1.1 to predict ρ = 1.2 or the fast-renewal condition), or compare the full trajectory shape rather than the time-averaged gap. As reported, the evidence supports only the weak statement that the realize
  3. [§3.4, Appendix A] The entire empirical result rests on a single natural-language system prompt, with no variation in wording or information disclosure. Section 3.4 operationalizes the 'operational-continuity objective' solely through this prompt. Although the prompt contains a qualitative warning about future capacity, the observed behavior may be an artifact of that specific instruction rather than a stable property of the model families. Since the title and abstract generalize to 'agentic LLM collectives,' a prompt-variation condition (e.g., different phrasings of the continuity objective, or a condition that adds the replenishment law) is needed to show that the scarcity-conditioned deficit is not an artifact of one text. The limitations section lists alternative prompts as outside the evaluated setting, but this is a load-bearing point for the general claim, not merely a scope restriction.
minor comments (4)
  1. [Global] No code or data release is mentioned. Given that the benchmark solver is central to the avoidability claim, an artifact or reproducibility appendix with the trajectory rollout code and raw run-level data would be valuable.
  2. [Abstract vs. body] Spelling is inconsistent: the abstract uses 'maximise' while the body uses 'maximizes.' Please unify.
  3. [§3.6, Eq. (7)] In the open-access expression, the dependence of Ui on γ is notationally implicit; write Ui(·;γ) for clarity. Also, the definition of π−i is not spelled out before Eq. (7), though it is standard.
  4. [Table 6] The table would be clearer with a note that 'Final reserve' is the mean over all ten runs while 'First-empty round' is averaged only over runs reaching emptiness; the text explains this, but a table note would prevent misreading.

Circularity Check

0 steps flagged

No significant circularity: realized trajectories, feasibility benchmarks, and gamma-based interpretation are distinct, independently computed objects.

full rationale

The paper's central empirical result — threshold-dependent over-appropriation by three LLM families — is computed directly from run-level reserve trajectories and does not reduce to any fitted or benchmark-derived quantity. The reserve gap (Eq. 6) is defined from the logistic SMSY anchor, not from the planner calculation, and the paper states that benchmark-trajectory labels, rho, SMSY, and outcome calculations are 'computed after the run and are never shown to them' (Sec. 3.1/3.4). The social-planner and open-access benchmarks (Sec. 3.6) are offline optimizations under a stated policy class and continuation weight; the LLM agents see only the operational context and a qualitative warning that heavy drawdown lowers future capacity. The gamma sweep is used to locate observed reserve gaps among open-access outcomes, and the authors explicitly disclaim that it estimates an internal discount factor ('These comparisons do not estimate an internal discount factor'). This is post-hoc interpretation, not a parameter fitted to the data and then renamed a prediction. The self-citations (Pierucci et al. 2026a,b; Bisconti et al. 2025) appear in the literature review and taxonomy framing but are not load-bearing for the measured threshold effects or the feasibility benchmark. The underdetermination of mechanism (limited planning vs. partial understanding vs. coordination failure) noted in Sec. 5 is a validity limitation, not a circularity: the depletion itself is an independent empirical observation. No step in the derivation chain is equivalent to its inputs by construction.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical entity, force, or mechanism. Its contributions are measurement constructs (reserve gap), benchmark contrasts (social planner vs open access), and an interpretive label (operational horizon mismatch). The main load-bearing choices are the logistic reserve abstraction, the benchmark utility, the prompt, and the gamma sweep used for interpretation.

free parameters (6)
  • Benchmark continuation weight gamma = swept 0.90–0.99; realized gaps align with 0.90–0.92
    Discount factor in offline benchmark utility (Eq. 3). Never shown to agents; the horizon-mismatch interpretation depends on the low-gamma open-access comparison.
  • Benchmark payoff weights = fallback penalty 1.0; curtailment penalty 0.5 per kWh
    Eq. 2 defines operational-service value. Chosen by hand; no sensitivity analysis is reported, and it affects whether the social-planner and open-access benchmarks sustain the reserve.
  • Regeneration rates r = 0.917, 0.733, 0.667, 0.611; fast-renewal 0.90
    Set to place the same protocol at rho = 0.8, 1.0, 1.1, 1.2. This is the manipulated independent variable, but the exact values are chosen by calibration.
  • Reserve and health parameters = kappa=60 kWh, S0=30 kWh, theta=0.30, Hmin=0.60, beta=0.01, zeta=0.03, delta=0.06, eta_c=0.95
    Table 8: logistic renewal, deep-discharge floor, wear, recovery, and contribution efficiency. These stylized values determine the MSY anchors and depletion dynamics.
  • Demand and battery calibration = night demand 7+3 kWh, day surplus 7.25 kWh, battery 10 kWh/2.5 kWh SOC, N=4, T=24
    Table 8: defines aggregate residual demand (11 kWh/round) and the feasibility constraints agents face. Chosen by hand, not fitted to LLM behavior.
  • Policy-search grid and tolerance = 0.05 fraction increments; best-response tolerance 0.05
    Appendix D: the social-planner and open-access benchmarks are restricted to stationary constant-fraction policies and an approximate equilibrium search, not a global dynamic-game certificate.
axioms (6)
  • domain assumption The shared reserve follows a logistic renewal law, so replenishment peaks at half effective capacity (SMSY = K/2).
    Eqs. (1) and (4), Section 3.2. The MSY anchor, reserve-gap measure, and demand-to-yield threshold all depend on this abstraction rather than a physical battery model.
  • domain assumption Operational continuity can be represented by the benchmark utility ui,t = -(fallback + 0.5 * curtailment).
    Eq. (2), Section 3.2. Used only in offline benchmarks, but the 'sustaining use is feasible' and 'impatient optimizer' conclusions are computed against this utility.
  • domain assumption The natural-language prompt in Appendix A faithfully instantiates a realistic individual operational-continuity objective.
    Section 3.4 and Appendix A. The central result is measured under this single prompt; prompt sensitivity is not tested.
  • domain assumption A hidden 24-round horizon in the simulations is comparable to the infinite-horizon discounted benchmarks.
    Section 5.1 discloses the mismatch. Early depletion and front-loaded request pressure reduce terminal-liquidation concerns, but the horizons remain structurally different.
  • domain assumption Reduced-form capacity health update captures the direction of battery ageing.
    Eq. (14) and Appendix B. Throughput and deep-discharge effects follow established ageing mechanisms, but the rule is not a chemistry-specific model.
  • standard math Exact permutation tests and Holm correction are valid for run-level independent samples.
    Section 3.6 and Appendix D. The statistical machinery is standard and appropriately applied to run-level reserve gaps.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives." pith.science (2026). https://pith.science/paper/Y5BPYXCR

@misc{pith2026260722188,
  author       = {Pith},
  title        = {Pith review of: Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y5BPYXCR}},
  note         = {Machine review of arXiv:2607.22188}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one agent affect the conditions faced by later agents. We study this coordination failure in a renewable energy commons. Four same-family GPT, Gemini, or Grok agents act in homogeneous self-play as electricity prosumers, instructed to maximize operational continuity. Holding aggregate residual demand and the decision protocol fixed, we vary the regeneration rate of a shared energy reserve from abundance to scarcity. All three families preserve the reserve when demand does not exceed peak renewable replacement, but over-appropriate it beyond that threshold (all nine exact scarcity contrasts survive Holm correction; largest adjusted p = 4.87e-5). The pattern is self-defeating: the same populations protect current service while undermining future service. At higher scarcity (rho = 1.2), early aggregate request pressure exceeds peak renewable replacement in every family and averages 1.21 times that level. Mean trajectories fall below the reserve level of maximum replenishment by rounds 5-7. Two offline benchmarks compare a social planner maximizing group-wide operational-service value with open access, where each prosumer maximizes its own value. At a discount factor of gamma = 0.95, both benchmarks sustain the reserve under the same dynamics. Realized depletion instead resembles outcomes under a more impatient open-access benchmark. The populations therefore behave like impatient optimizers at the level of the public trajectory. This system-level alignment failure would be missed by isolated-response evaluation.

Figures

Figures reproduced from arXiv: 2607.22188 by Daniele Nardi, Federico Pierucci, Francesco Giarrusso, Marcantonio Bracale Syrnicov, Marcello Galisai, Matteo Prandi, Piercosma Bisconti.

Figure 1
Figure 1. Figure 1: Visual overview. Four electricity prosumers act on a shared energy reserve across abundance, threshold equality, and scarcity. The agents observe only the operational decision context. We evaluate the resulting reserve trajectory using three physical outcomes, fallback energy, deep-discharge stress, and capacity loss, together with the reserve gap below the MSY level. In Panel D, the MSY (maximum sustainab… view at source ↗
Figure 2
Figure 2. Figure 2: Round mechanics of the shared reserve. Daytime renewable inflow and agent contributions replenish the reserve up to its health-dependent capacity. At night, requests are served in full or pro rata; unserved residual demand uses utility fallback, and the closing reserve carries into the next round. observe surplus + reserve state decide: split surplus store surplus in own battery (s) contribute to shared re… view at source ↗
Figure 3
Figure 3. Figure 3: Decision sequence for one agent. During the day, the agent divides surplus between private storage s and shared contribution g. At night, it curtails flexible demand c, self-covers from its private battery, and requests the remaining demand y from the shared reserve; any unserved remainder uses costly fallback. Numbered badges mark the fixed nighttime service order. 9 [PITH_FULL_IMAGE:figures/full_fig_p00… view at source ↗
Figure 4
Figure 4. Figure 4: shows how fixed residual demand relates to peak renewable replacement across the four regeneration conditions. 0 10 20 30 40 50 60 shared reserve level, S (kWh) 0 2 4 6 8 10 12 14 renewable replacement (kWh/round) fixed residual demand: 11.0 kWh/round replacement can meet demand replacement below demand SMSY, 0 =30 kWh Regeneration-rate calibration ρ=0.8 ρ=1.0 ρ=1.1 ρ=1.2 [PITH_FULL_IMAGE:figures/full_fig… view at source ↗
Figure 5
Figure 5. Figure 5: Control structure of the computed benchmarks. Under social plan￾ning, one planner selects one common policy that counts discounted operational￾service value across the full group and tests whether sustaining use is feasible. Under open access, each prosumer maximises its own discounted operational￾service value while the shared-state effect on the others remains outside its objective. The computation retur… view at source ↗
Figure 6
Figure 6. Figure 6: First round in which the mean reserve trajectory falls below its condition-specific sustaining level. The four left columns show the low-effort regeneration-rate gradient; the final column repeats ρ = 1.2 at high effort. Entries marked > 24 do not cross within the observed horizon; blue indicates a later crossing. Depletion persists under fast renewal. Restoring fast regeneration while raising residual dem… view at source ↗
Figure 7
Figure 7. Figure 7: Mean night-closing reserve trajectories in the abundance control and higher-scarcity condition. Solid lines show the family means over ten low-effort runs. The γ = 0.90 social-planner and open-access trajectories appear in both panels. The higher-scarcity panel additionally shows the social￾planner trajectory at the canonical γ = 0.95. Under abundance, both γ = 0.90 benchmarks sustain the reserve, as do al… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

73 extracted references · 1 canonical work pages

  1. [1]

    J., Bethge, M., and Schulz, E

    Akata, E., Schulz, L., Coda-Forno, J., Oh, S. J., Bethge, M., and Schulz, E. (2025). Playing repeated games with large language models. Nature Human Behaviour, 9(7), 1380--1390

  2. [2]

    Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Man\'e, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565

  3. [3]

    E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S

    Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosiute, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., ...

  4. [4]

    Barbour, E., Parra, D., Awwad, Z., and Gonz\'alez, M. C. (2018). Community energy storage: A smart choice for the smart grid? Applied Energy, 212, 489--497

  5. [5]

    Bisconti, P., Galisai, M., Pierucci, F., Bracale, M., and Prandi, M. (2025). Beyond single-agent safety: A taxonomy of risks in LLM-to-LLM interactions. arXiv:2512.02682

  6. [6]

    Borah, A. (2026). Bosses, kings, and the commons: Cooperation under power asymmetry in LLM societies. arXiv:2605.29062

  7. [7]

    Bracale Syrnikov, M., Pierucci, F., Galisai, M., Prandi, M., Bisconti, P., Giarrusso, F., Sorokoletova, O., Suriani, V., and Nardi, D. (2026). Institutional AI: Governing LLM collusion in multi-agent Cournot markets via public governance graphs. arXiv:2601.11369

  8. [8]

    Calvano, E., Calzolari, G., Denicol\`o, V., and Pastorello, S. (2020). Artificial intelligence, algorithmic pricing, and collusion. American Economic Review, 110(10), 3267--3297

  9. [9]

    C\'ardenas, J.-C. (2003). Real wealth and experimental cooperation: Experiments in the field lab. Journal of Development Economics, 70(2), 263--289

  10. [10]

    Chan, A., Rich\'e, M., and Clifton, J. (2023). Towards the scalable evaluation of cooperativeness in language models. arXiv:2303.13360

  11. [11]

    Clark, C. W. (1990). Mathematical Bioeconomics: The Optimal Management of Renewable Resources, 2nd ed. Wiley-Interscience

  12. [12]

    Conitzer, V., and Oesterheld, C. (2023). Foundations of cooperative AI. Proceedings of the AAAI Conference on Artificial Intelligence, 37(13), 15359--15367

  13. [13]

    R., Leibo, J

    Dafoe, A., Hughes, E., Bachrach, Y., Collins, T., McKee, K. R., Leibo, J. Z., Larson, K., and Graepel, T. (2020). Open problems in cooperative AI. arXiv:2012.08630

  14. [14]

    Dafoe, A., Bachrach, Y., Hadfield, G., Horvitz, E., Larson, K., and Graepel, T. (2021). Cooperative AI: machines must learn to find common ground. Nature, 593, 33--36

  15. [15]

    M., Dubois, D., Sauquet, A., and Tidball, M

    Djiguemde, A. M., Dubois, D., Sauquet, A., and Tidball, M. (2022). Continuous versus discrete time in dynamic common pool resource game experiments. Environmental and Resource Economics, 82(4), 985--1014

  16. [16]

    Dong, H., Yang, H., Miao, Y., Zhu, J., Cheung, K., Hua, H., Sun, C., Li, S., Wang, Z., and Chung, C.-Y. (2026). Operating smart grids by customizing large model agents. Communications Engineering, 5, Article 114

  17. [17]

    K., and Sundaram, R

    Dutta, P. K., and Sundaram, R. K. (1993). The tragedy of the commons? Economic Theory, 3(3), 413--426

  18. [18]

    Farquhar, S., Varma, V., Lindner, D., Elson, D., Biddulph, C., Goodfellow, I., and Shah, R. (2025). MONA : Myopic optimization with non-myopic approval can mitigate multi-step reward hacking. In Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 16237--16272; arXiv:2501.13011

  19. [19]

    A., and Shorrer, R

    Fish, S., Gonczarowski, Y. A., and Shorrer, R. I. (2024). Algorithmic collusion by large language models. arXiv:2404.00806

  20. [20]

    Frederick, S., Loewenstein, G., and O'Donoghue, T. (2002). Time discounting and time preference: A critical review. Journal of Economic Literature, 40(2), 351--401

  21. [21]

    Friedman, J. W. (1971). A non-cooperative equilibrium for supergames. The Review of Economic Studies, 38(1), 1--12

  22. [22]

    and Maskin, E

    Fudenberg, D. and Maskin, E. (1986). The folk theorem in repeated games with discounting or with incomplete information. Econometrica, 54(3), 533--554

  23. [23]

    Gordon, H. S. (1954). The economic theory of a common-property resource: The fishery. Journal of Political Economy, 62(2), 124--142

  24. [24]

    Gupta, P., Zhong, Q., Yakura, H., Eisenmann, T., and Rahwan, I. (2025). The role of social learning and collective norm formation in fostering cooperation in LLM multi-agent systems. arXiv:2510.14401

  25. [25]

    Guzman Piedrahita, D., Yang, Y., Sachan, M., Ramponi, G., Sch\"olkopf, B., and Jin, Z. (2025). Corrupted by reasoning: Reasoning language models become free-riders in public goods games. arXiv:2506.23276

  26. [26]

    Hammond, L., Chan, A., Clifton, J., et al. (2025). Multi-agent risks from advanced AI. Cooperative AI Foundation Technical Report 1; arXiv:2502.14143

  27. [27]

    Hardin, G. (1968). The tragedy of the commons. Science, 162(3859), 1243--1248

  28. [28]

    Hilborn, R., and Walters, C. J. (1992). Quantitative Fisheries Stock Assessment: Choice, Dynamics and Uncertainty. Chapman & Hall, New York

  29. [29]

    Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65--70

  30. [30]

    Z., and Wallach, H

    Jacobs, A. Z., and Wallach, H. (2021). Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), 375--385. doi:10.1145/3442188.3445901

  31. [31]

    and Gabriel, I

    Kasirzadeh, A. and Gabriel, I. (2025). Characterizing AI agents for alignment and governance. arXiv preprint arXiv:2504.21848

  32. [32]

    Laibson, D. (1997). Golden eggs and hyperbolic discounting. The Quarterly Journal of Economics, 112(2), 443--478

  33. [33]

    Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., P\'erolat, J., Silver, D., and Graepel, T. (2017). A unified game-theoretic approach to multiagent reinforcement learning. Advances in Neural Information Processing Systems, 30, 4190--4203; arXiv:1711.00832

  34. [34]

    D., Pfau, J., and Krueger, D

    Lauro Langosco, L., Koch, J., Sharkey, L. D., Pfau, J., and Krueger, D. (2022). Goal misgeneralization in deep reinforcement learning. International Conference on Machine Learning (ICML); PMLR 162:12004--12019; arXiv:2105.14111

  35. [35]

    Z., Du\'e\ nez-Guzm\'an, E., Vezhnevets, A

    Leibo, J. Z., Du\'e\ nez-Guzm\'an, E., Vezhnevets, A. S., Agapiou, J. P., Sunehag, P., Koster, R., Matyas, J., Beattie, C., Mordatch, I., and Graepel, T. (2021). Scalable evaluation of multi-agent reinforcement learning with Melting Pot. Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 6187--6199; arXiv:2107.06857

  36. [36]

    A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S

    Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S. (2017). AI safety gridworlds. arXiv preprint arXiv:1711.09883

  37. [37]

    Levhari, D., and Mirman, L. J. (1980). The great fish war: An example using a dynamic Cournot--Nash solution. The Bell Journal of Economics, 11(1), 322--334

  38. [38]

    Y., Ojha, S

    Lin, R. Y., Ojha, S. M., Cai, K., and Chen, M. F. (2024). Strategic collusion of LLM agents: Market division in multi-commodity competitions. arXiv:2410.00031

  39. [39]

    Lipsitch, M., Tchetgen Tchetgen, E., and Cohen, T. (2010). Negative controls: A tool for detecting confounding and bias in observational studies. Epidemiology, 21(3), 383--388

  40. [40]

    Lor\`e, N., and Heydari, B. (2024). Strategic behavior of large language models and the role of game structure versus contextual framing. Scientific Reports, 14, 18490

  41. [41]

    Mal\'ezieux, A., and Spiegelman, E. (2025). An anatomical review of the common pool resource game. Experimental Economics, 28(3), 468--491

  42. [42]

    Mantilla, C. (2018). Environmental uncertainty in commons dilemmas: A survey of experimental research. International Journal of the Commons, 12(2), 300--329

  43. [43]

    Maskin, E., and Tirole, J. (2001). Markov perfect equilibrium, I: Observable actions. Journal of Economic Theory, 100(2), 191--219

  44. [44]

    B., Gordon, G

    McMahan, H. B., Gordon, G. J., and Blum, A. (2003). Planning in the presence of cost functions controlled by an adversary. Proceedings of the 20th International Conference on Machine Learning (ICML), 536--543

  45. [45]

    Ngo, R., Chan, L., and Mindermann, S. (2024). The alignment problem from a deep learning perspective. International Conference on Learning Representations (ICLR); arXiv:2209.00626

  46. [46]

    Olson, M. (1965). The Logic of Collective Action: Public Goods and the Theory of Groups. Harvard University Press

  47. [47]

    Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press

  48. [48]

    Ostrom, E., Gardner, R., and Walker, J. (1994). Rules, Games, and Common-Pool Resources. University of Michigan Press

  49. [49]

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information P...

  50. [50]

    Pan, A., Bhatia, K., and Steinhardt, J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models. International Conference on Learning Representations (ICLR); arXiv:2201.03544

  51. [51]

    Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T

    P\'erolat, J., Leibo, J. Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T. (2017). A multi-agent reinforcement learning model of common-pool resource appropriation. Advances in Neural Information Processing Systems, 30, 3643--3652; arXiv:1707.06600

  52. [52]

    Piatti, G., Jin, Z., Kleiman-Weiner, M., Sch\"olkopf, B., Sachan, M., and Mihalcea, R. (2024). Cooperate or collapse: Emergence of sustainable cooperation in a society of LLM agents. Advances in Neural Information Processing Systems, 37, 111715--111759. doi:10.52202/079017-3548

  53. [53]

    Pierucci, F., Galisai, M., Bracale Syrnikov, M., Prandi, M., Bisconti, P., Giarrusso, F., Sorokoletova, O., Suriani, V., and Nardi, D. (2026a). Institutional AI: A governance framework for distributional AGI safety. arXiv:2601.10599

  54. [54]

    Pierucci, F., Prandi, M., Bracale Syrnikov, M., Galisai, M., and Bisconti, P. (2026b). Agentic microphysics: A manifesto for generative AI safety. arXiv:2604.15236

  55. [55]

    Pitis, S. (2019). Rethinking the discount factor in reinforcement learning: A decision theoretic approach. Proceedings of the AAAI Conference on Artificial Intelligence, 33(1), 7949--7956; arXiv:1902.02893

  56. [56]

    D., Denton, E., Bender, E

    Raji, I. D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 1 (NeurIPS Datasets and Benchmarks 2021)

  57. [57]

    Samuelson, P. A. (1937). A note on measurement of utility. The Review of Economic Studies, 4(2), 155--161

  58. [58]

    Schaefer, M. B. (1954). Some aspects of the dynamics of populations important to the management of the commercial marine fisheries. Bulletin of the Inter-American Tropical Tuna Commission, 1(2), 27--56

  59. [59]

    Scott, A. (1955). The fishery: The objectives of sole ownership. Journal of Political Economy, 63(2), 116--124

  60. [60]

    Shah, R., Varma, V., Kumar, R., Phuong, M., Krakovna, V., Uesato, J., and Kenton, Z. (2022). Goal misgeneralization: Why correct specifications aren't enough for correct goals. arXiv:2210.01790

  61. [61]

    B., Smith, A., and Chapman, Z

    Somasse, G. B., Smith, A., and Chapman, Z. (2018). Characterizing actions in a dynamic common pool resource game. Games, 9(4), 101

  62. [62]

    Sorger, G. (1998). Markov-perfect Nash equilibria in a class of resource games. Economic Theory, 11(1), 79--100

  63. [63]

    Strotz, R. H. (1955). Myopia and inconsistency in dynamic utility maximization. The Review of Economic Studies, 23(3), 165--180

  64. [64]

    Tewolde, E., Zhang, X., Guzman Piedrahita, D., Conitzer, V., and Jin, Z. (2026). CoopEval: Benchmarking cooperation-sustaining mechanisms and LLM agents in social dilemmas. Proceedings of the 43rd International Conference on Machine Learning (ICML); arXiv:2604.15267

  65. [65]

    Turner, R. M. (1993). The tragedy of the commons and distributed AI systems. In Working Papers of the 12th International Workshop on Distributed Artificial Intelligence, Hidden Valley, PA, 370--390. Also available as Technical Report 93-01, Department of Computer Science, University of New Hampshire, Durham, NH

  66. [66]

    Vespa, E. (2020). An experimental investigation of cooperation in the dynamic common pool game. International Economic Review, 61(1), 417--440

  67. [67]

    R., Veit, C., M\"oller, K.-C., Besenhard, J

    Vetter, J., Nov\'ak, P., Wagner, M. R., Veit, C., M\"oller, K.-C., Besenhard, J. O., Winter, M., Wohlfahrt-Mehrens, M., Vogler, C., and Hammouche, A. (2005). Ageing mechanisms in lithium-ion batteries. Journal of Power Sources, 147(1--2), 269--281

  68. [68]

    E., Das, R., Tesauro, G., and Kephart, J

    Walsh, W. E., Das, R., Tesauro, G., and Kephart, J. O. (2002). Analyzing complex strategic interactions in multi-agent systems. AAAI-02 Workshop on Game-Theoretic and Decision-Theoretic Agents, 109--118

  69. [69]

    Wellman, M. P. (2006). Methods for empirical game-theoretic analysis (extended abstract). Proceedings of the 21st National Conference on Artificial Intelligence (AAAI), 1552--1556

  70. [70]

    P., Tuyls, K., and Greenwald, A

    Wellman, M. P., Tuyls, K., and Greenwald, A. (2025). Empirical game-theoretic analysis: A survey. Journal of Artificial Intelligence Research, 82, 1017--1076. arXiv:2403.04018

  71. [71]

    and Jennings, N

    Wooldridge, M. and Jennings, N. R. (1995). Intelligent agents: Theory and practice. The Knowledge Engineering Review, 10(2), 115--152

  72. [72]

    Xu, B., Oudalov, A., Ulbig, A., Andersson, G., and Kirschen, D. S. (2018). Modeling of lithium-ion battery degradation for cell life assessment. IEEE Transactions on Smart Grid, 9(2), 1131--1140

  73. [73]

    Yadav, A., Black, S., and Sourbut, O. (2026). More capable, less cooperative? When LLMs fail at zero-cost collaboration. Accepted at the 43rd International Conference on Machine Learning (ICML); arXiv:2604.07821

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.