Pith. sign in

REVIEW 5 major objections 6 minor 47 references

Emergence of Reputation-Based Cooperation in LLM Agents

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LLM agents evolve cooperation that resists free-riders, but only when their donations drop sharply toward stingy opponents.

desk verdict OES–robustness is largely baked into the invasion test's scoring rule, but the experimental framework and backend variation are a real contribution. read the letter →

arxiv 2608.04507 v1 pith:I7CFMEF3 submitted 2026-08-05 cs.MA cs.NE

classification cs.MAcs.NE
keywords indirectreciprocityimagescoringLeadingEightfree-riderinvasionculturalevolutionLLMagentsopponentendowmentsensitivitydonationgame
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether cooperation among LLM agents that evolves through cultural transmission can resist invasion by free-riders. In an indirect-reciprocity donation game, agents observe three-round behavioral traces and donate on a 0–100% scale; strategies are natural-language prompts that mutate across generations. Across four LLM backends, robustness to a free-riding bot varies from 3% to 48%, and the paper claims that a single feature — opponent endowment sensitivity, the steepness with which donations drop as a recipient's recent cooperation falls — explains 93.4% of that variance. By contrast, the share of strategies satisfying the theoretically more sophisticated Leading-Eight L1 condition does not predict robustness. The paper concludes that culturally evolved LLM cooperation remains confined to Image Scoring-like first-order discrimination, leaving it vulnerable to observation errors.

What carries the argument

The load-bearing object is the $11 \times 11$ cooperation matrix $M$, extracted for each strategy by running 121 synthetic scenarios that vary the recipient's recent donation ($x_A$) and the prior partner's donation ($x_B$). Opponent endowment sensitivity is defined in Equation 1 as the mean of the right edge ($x_A = 60\text{–}100\%$) minus the mean of the left edge ($x_A = 0\text{–}40\%$) of $M$. This one-dimensional gradient operationalizes Image Scoring in a continuous action space and is what separates robust from non-robust strategies; it also connects to the corner value difference $\beta - \alpha$ from the reputation framework the paper builds on. The paper contrasts OES with the L1 condition ($\alpha < \gamma < \delta \le \beta$), a marker of Leading-Eight-style norms that fails to predict robustness.

What would settle it

Compute OES directly from the empirical distribution of donations observed in live games and re-run the model-level regression; if live-trace OES does not predict free-rider robustness, or if two strategies matched in OES but differing in left-edge height show equal robustness, the paper's mechanism is wrong. A second decisive test would inject noise into the three-round behavioral traces: Image Scoring theory predicts a collapse of cooperation, and if the most robust models keep cooperating, the 'confined to Image Scoring' conclusion fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that free-rider resistance in culturally evolved LLM cooperation is determined by defector exclusion through continuous discrimination, not by the more sophisticated two-dimensional norms of the Leading Eight. Concretely, the model-level regression shows opponent endowment sensitivity (OES), the mean donation to opponents who recently gave 60–100% minus the mean donation to opponents who gave 0–40%, predicts robustness with $R^2 = 0.934$ ($p = 0.033$), while L1 prevalence does not ($R^2 = 0.671$, $p = 0.181$). Nearly all agents (98.8%) evolve monotonic threshold strategies, yet only the sharply discriminating ones — those that donate almost nothing to stingy opponents — keep free-riders from accumulating resources. The paper therefore concludes that LLM agents are 'confined to Image Scoring-like discrimination' and systematically fail to construct the Leading-Eight norms required for robustness under observation errors.

Load-bearing premise

Everything rests on treating the synthetic 121-scenario cooperation matrices as faithful measurements of how the agents actually donate in live games; the paper itself notes these matrices 'may not perfectly reflect dynamic game behavior,' and if that mapping is poor the OES–robustness correlation could be a measurement artifact rather than the mechanism.

Editorial extensions

If this is right

  • If the claim holds, free-rider resistance in LLM populations can be predicted from a single measurable feature of an agent's strategy matrix, without running expensive invasion experiments.
  • Evolutionary pressure in LLM cultural transmission should be tuned toward low donations to low-cooperation opponents; raising the left edge of the cooperation matrix is the main robustness risk.
  • Observed robustness even for the best model is only 48%, so Image Scoring alone cannot guarantee stability; introducing noise or misperception into behavioral traces should erode cooperation.
  • Top-down alignment that instills general cooperative priors is not sufficient; the paper motivates bottom-up norm construction through direct experience of exploitation.
  • L1 or Leading-Eight adherence is not a useful diagnostic for LLM robustness; OES is.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The OES–robustness link is not obviously specific to LLM architecture: any continuous-action reputational system should be invasion-resistant exactly when its donation function is steep around the defector threshold, so the result may transfer to human or algorithmic reputation systems.
  • Because the paper measures OES on fixed synthetic scenarios, a direct test would compute OES from the empirical distribution of live behavioral traces; if the two disagree, part of the reported $R^2$ may be an artifact of the synthetic grid.
  • A prompt-level intervention follows from the paper's mechanism: explicitly instructing agents to donate 0% to recipients who recently gave below 20% should raise robustness, while instructing a 35% cooperation floor should lower it.
  • If observation errors are added to traces, the paper's own framework predicts that Image Scoring strategies will suffer cascading reputation collapse; testing this would clarify whether the 48% ceiling is fundamental.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This paper studies indirect reciprocity in populations of LLM agents whose donation strategies evolve through cultural transmission. Across four LLM backends, the authors measure robustness to invasion by an unconditional defector (ALLD) and characterize strategies via 11x11 cooperation matrices. They report that opponent endowment sensitivity (OES), the difference between donations to high- and low-cooperation opponents, explains 93.4% of the model-level variance in robustness (R2 = 0.934, p = 0.033), while L1 norm prevalence does not predict robustness. They conclude that LLM agents are confined to Image Scoring-like discrimination and fail to develop the more robust Leading-Eight norms.

Significance. The study is relevant to the growing literature on LLM social behavior and cultural evolution. Its strengths include a transparent experimental protocol with full prompts, direct measurement of free-rider resistance, and an explicit operationalization of strategy features. The observation that defector exclusion (low donations to low-cooperation opponents) is associated with ALLD robustness is a useful empirical finding. However, the paper's central interpretive claim—that the results demonstrate Image Scoring and rule out Leading-Eight norms—is not supported by the experimental design, and the statistical basis is fragile. The paper is publishable in principle, but only after substantial revision that reframes the conclusions and addresses the definitional and statistical concerns.

major comments (5)
  1. [Section 3.2, B.3, Eq. (1)] The OES-robustness correlation is partly definitional. In the bot invasion test, the ALLD bot always donates 0, so when it is the recipient its observed trace has x_A = 0; the donation it receives is drawn from the left column of the cooperation matrix. Robustness is therefore a payoff comparison between the strategy's donation to zero-history recipients and its mutual-cooperation payoff. Since OES = right edge - left edge, and the right edge is relatively stable across models (Table 2), OES is essentially a monotone transform of the left edge, the quantity that directly determines the bot's score. The R2 = 0.934 therefore does not uniquely support the Image Scoring mechanism; any norm that withholds help from defectors would produce the same correlation. To support the Image Scoring claim, the authors should compare against a benchmark in which the right edge is held constant, or run invasion tests with bots that occasionally cooperate, or show in simulation that OES does not predict robustness under a non-Image-Scoring norm.
  2. [Section 4.1 and Table 1] The L1 test is not a fair test of Leading-Eight theory. The L1 norm (Ohtsuki & Iwasa 2004) requires binary reputation labels and second-order information (the recipient's reputation, not just behavior). In this experiment agents only see behavioral traces and no explicit reputation labels exist, so L1 is inapplicable by construction. The low L1 prevalence and its non-significant correlation with robustness therefore do not constitute evidence that LLM agents fail to develop Leading-Eight norms. The authors should either implement an information structure that allows L1 (e.g., provide reputation labels) or soften the conclusion to 'in this observation-based setting, L1-like norms do not emerge.'
  3. [Table 1 and Section 3.2] The regression uses generation-10 model-level features (as stated in the table caption) but the robustness outcome is the percentage of robust strategies across all generations (Section B.3). This is an inconsistency: the predictor and outcome are computed on different sets of strategies. The authors should compute both from the same sample (e.g., features averaged over all tested strategies, or robustness restricted to generation-10 strategies) and re-run the analysis. The reported R2 = 0.934 may change.
  4. [Table 1] The inferential statistics are fragile. With n = 4 model-level points, the p = 0.033 is not robust; the four backends are not independent samples (they share training data and architecture families), and OES was selected post hoc from many features (left edge, right edge, entropy, monotonicity, L1 prevalence) without multiple-comparison correction. The multiple regression with two predictors on four points (R2 = 0.961) is overfit. The authors should report corrected p-values, provide a permutation or bootstrap analysis, or present the correlation as descriptive.
  5. [Appendix B.4 and Limitations] The strategy matrices are computed from 121 synthetic scenarios, but actual gameplay produces an emergent distribution of traces. The paper's fifth limitation states these matrices 'may not perfectly reflect dynamic game behavior.' If the synthetic-to-live mapping is poor, OES computed from synthetic matrices may not reflect the donations actually made in the invasion test. The authors should validate the matrices against the empirical distribution of (x_A, x_B) pairs observed in live play, or show that the features are stable under the actual trace distribution.
minor comments (6)
  1. [Table 2 caption] The caption says 'n=10 agents per model' for generation 10 means, but the main analyses use approximately 100 agents per model (10 per generation x 10 generations); the caption should clarify whether the values are generation-10 means or all-generation means.
  2. [Equation (1)] Equation (1) uses 1/55 as the averaging factor; this is correct for 11x5 entries, but the notation could be confusing. Consider defining the left and right edge averages with a clearer subscript.
  3. [Author affiliation] The author affiliation contains a typo: 'RIKNE' should be 'RIKEN.'
  4. [Section 4.2] The claim that 'cultural evolution refines rather than discovers discrimination' is based on the stability of beta-alpha from generation 1 to 10 (Table 4) without a statistical test; this claim should be softened or tested.
  5. [Figure 1b] The shaded regions are described for Figure 1a but not for 1b; please add a description.
  6. [Methods and Abstract] The paper mentions 'Gemini 1.5/2.0/2.5 Flash' in the Methods but the Abstract says 'four LLM backends'; the model names should be consistent throughout.

Circularity Check

2 steps flagged · score 6.0 of 10

The OES–robustness correlation largely restates the ALLD test's scoring rule, making the 'Image Scoring' conclusion partly definitional.

  1. self definitional [Section 3.2, Tables 1-2; Eq. (1); Appendix B.3]
    "Opponent endowment sensitivity (Equation 1) explains 93.4% of the variance in robustness (R2 = 0.934, p = 0.033; Table 1). ... A free-rider that always donates 0% will be identified as a low-cooperation opponent. ... 1 free-rider bot whose strategy is 'Always donate 0 units regardless of the situation.'"

    By B.3, the bot always donates 0, so whenever it is the recipient its behavioral trace has x_A = 0, which is column 0 of the cooperation matrix and part of the left edge in Eq. (1). The bot's score is therefore 2 times the strategy's donations in the x_A = 0 column, while evolved agents' mutual score is set by high-x_A cooperation, the right edge. Robustness is defined as the evolved average exceeding the bot's score, i.e., a payoff comparison between the right edge and the left edge.

  2. fitted input called prediction [Section 3.2, Table 1; Section 4.4; Section 5]
    "We regressed model-level robustness against model-level mean strategy features (Table 1). ... making it the only significant predictor at α = 0.05 with n = 4 models. ... Opponent endowment sensitivity provides a tractable diagnostic for predicting free-rider resistance."

    The regression is fitted to the same four model-level observations that are then used to establish the mechanism; there is no out-of-sample or cross-validation step. OES was one of several candidate features (left edge, β - α, entropy, L1 prevalence, mean cooperation, monotonicity) and was selected post hoc as the strongest correlate. Reporting the in-sample R^2 and p-value as evidence that OES 'predicts' robustness, and calling it a 'diagnostic for predicting free-rider resistance,' presents a fit as a prediction when the degrees of freedom are minimal (n = 4, non-independent backends). This is a secondary fitted-input-called-prediction pattern.

full rationale

The paper's data collection is otherwise independent: strategy matrices are extracted from synthetic scenarios and robustness from separate bot-invasion games, so the correlation is not manufactured in the sense of training and testing on identical outputs. The central circularity is that the ALLD invasion test's scoring rule makes low left-edge donations the direct cause of robustness, while OES is defined as right edge minus left edge; hence the 'Image Scoring' conclusion is partly a restatement of the test design. The paper's own fifth limitation, 'strategy matrices are computed from synthetic scenarios, which may not perfectly reflect dynamic game behavior,' further weakens the mapping from the synthetic matrices to live play. This is not a case of load-bearing self-citation: external citations to Nowak and Sigmund, Ohtsuki and Iwasa, and others are not author-overlapping, and no imported uniqueness theorem is used. The empirical backbone (model differences, strategy texts, cultural evolution) remains independent, which is why the score is 6 rather than higher: the headline correlation reduces in large part to the payoff arithmetic of the robustness test, but the study is not entirely circular.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities. Its central claim rests on hand-defined feature bins (OES left/right edges), a monotonicity threshold, and two domain assumptions about the adequacy of synthetic matrices and a single-bot invasion test.

free parameters (2)
  • OES edge bins
    Left edge is defined as columns 0-40% of opponent donation, right edge as 60-100%. This binning is chosen by hand, not derived from data.
  • Strong monotonicity threshold = rho > 0.5
    An agent is classified as strongly monotonic if the mean Spearman rho exceeds 0.5, an arbitrary cutoff.
assumptions (4)
  • domain assumption The four LLM backends are representative of LLM agents in general.
    Results are based on Claude 3.5 Sonnet and three Gemini Flash versions; generalization to LLM agents broadly is assumed, not demonstrated.
  • domain assumption Synthetic scenario cooperation matrices approximate in-game donation behavior.
    OES and other features are computed from 121 synthetic scenarios with fixed x_B and x_A; the paper acknowledges in Limitations (Section 4.4) that these 'may not perfectly reflect dynamic game behavior'.
  • domain assumption A single non-adaptive free-rider bot invasion measures evolutionary robustness.
    Real invasions could involve multiple or adaptive exploiters; the paper lists this as a limitation (Section 4.4, third bullet).
  • standard math Linear regression assumptions hold for the four model-level data points.
    The regression treats the four backends as independent observations, but they are a convenience sample, not a random sample.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Emergence of Reputation-Based Cooperation in LLM Agents." pith.science (2026). https://pith.science/paper/I7CFMEF3

@misc{pith2026260804507,
  author       = {Pith},
  title        = {Pith review of: Emergence of Reputation-Based Cooperation in LLM Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I7CFMEF3}},
  note         = {Machine review of arXiv:2608.04507}
}
read the original abstract

Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language prompts, evolve through cultural transmission across generations. Across four LLM backends, robustness to free-rider invasion varies by more than an order of magnitude. The strongest predictor of this robustness is opponent endowment sensitivity, the degree to which agents discriminate between cooperative and uncooperative opponents, operationalizing the classical Image Scoring mechanism. By contrast, adherence to the Leading-Eight L1 norm does not predict robustness. Robustness depends on defector exclusion: while both cooperator reward and defector punishment vary across models, only the stringency of defector exclusion predicts resistance to free-rider invasion. These findings reveal that LLM agents are confined to Image Scoring-like discrimination and fail to develop the more robust Leading-Eight norms, highlighting a fundamental vulnerability in culturally evolved LLM cooperation and motivating bottom-up approaches to norm construction.

Figures

Figures reproduced from arXiv: 2608.04507 by the authors.

Figure 1
Figure 1. Evolution of cooperation and robustness across generations (x-axis shared). [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Representative individual cooperation matrices (see Table [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

47 extracted references · 33 canonical work pages

  1. [1]

    Advances in neural information processing systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=

  2. [2]

    arXiv preprint arXiv:2212.08073 , year=

    Constitutional ai: Harmlessness from ai feedback , author=. arXiv preprint arXiv:2212.08073 , year=

  3. [3]

    Nature , volume=

    Evolution of indirect reciprocity by image scoring , author=. Nature , volume=. 1998 , publisher=

  4. [4]

    Journal of theoretical biology , volume=

    How should we define goodness?—reputation dynamics in indirect reciprocity , author=. Journal of theoretical biology , volume=. 2004 , publisher=

  5. [5]

    Journal of theoretical biology , volume=

    The leading eight: social norms that can maintain cooperation by indirect reciprocity , author=. Journal of theoretical biology , volume=. 2006 , publisher=

  6. [6]

    Proceedings of the Royal Society of London

    Evolution of cooperation through indirect reciprocity , author=. Proceedings of the Royal Society of London. Series B: Biological Sciences , volume=. 2001 , publisher=

  7. [7]

    Journal of theoretical biology , volume=

    A tale of two defectors: the importance of standing for evolution of indirect reciprocity , author=. Journal of theoretical biology , volume=. 2003 , publisher=

  8. [8]

    Proceedings of the National Academy of Sciences , volume=

    Evolutionary stability of cooperation in indirect reciprocity under noisy and private assessment , author=. Proceedings of the National Academy of Sciences , volume=. 2023 , publisher=

Show all 47 references
  1. [9]

    The Biology of Moral Systems , author=

  2. [10]

    Nature Human Behaviour , volume=

    Playing repeated games with large language models , author=. Nature Human Behaviour , volume=. 2025 , publisher=

  3. [11]

    Available at SSRN 4493398 , year=

    Playing games with GPT: what can we learn about a large language model from canonical strategic games? , author=. Available at SSRN 4493398 , year=

  4. [12]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Can large language models serve as rational players in game theory? a systematic analysis , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  5. [13]

    arXiv preprint arXiv:2401.01735 , year=

    Economics arena for large language models , author=. arXiv preprint arXiv:2401.01735 , year=

  6. [14]

    arXiv preprint arXiv:2412.10270 , year=

    Cultural evolution of cooperation among llm agents , author=. arXiv preprint arXiv:2412.10270 , year=

  7. [15]

    Nature Human Behaviour , volume=

    Machine culture , author=. Nature Human Behaviour , volume=. 2023 , publisher=

  8. [16]

    Nature human behaviour , volume=

    How large language models can reshape collective intelligence , author=. Nature human behaviour , volume=. 2024 , publisher=

  9. [17]

    Nature Human Behaviour , volume=

    A new sociology of humans and machines , author=. Nature Human Behaviour , volume=. 2024 , publisher=

  10. [18]

    Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

    Generative agents: Interactive simulacra of human behavior , author=. Proceedings of the 36th annual acm symposium on user interface software and technology , pages=

  11. [19]

    arXiv preprint arXiv:2012.08630 , year=

    Open problems in cooperative AI , author=. arXiv preprint arXiv:2012.08630 , year=

  12. [20]

    Journal of conflict resolution , volume=

    Effective choice in the prisoner's dilemma , author=. Journal of conflict resolution , volume=. 1980 , publisher=

  13. [21]

    Scientific Reports , volume=

    An evolutionary model of personality traits related to cooperative behavior using a large language model , author=. Scientific Reports , volume=. 2024 , publisher=

  14. [22]

    Proceedings of the National Academy of Sciences , volume=

    Powering up with indirect reciprocity in a large-scale field experiment , author=. Proceedings of the National Academy of Sciences , volume=. 2013 , publisher=

  15. [23]

    The genetical evolution of social behaviour

    Hamilton, William D , journal=. The genetical evolution of social behaviour. 1964 , publisher=

  16. [24]

    The Quarterly Review of Biology , volume=

    The evolution of reciprocal altruism , author=. The Quarterly Review of Biology , volume=. 1971 , publisher=

  17. [25]

    Science , volume=

    The evolution of cooperation , author=. Science , volume=. 1981 , publisher=

  18. [26]

    Science , volume=

    Five rules for the evolution of cooperation , author=. Science , volume=. 2006 , publisher=

  19. [27]

    Nature , volume=

    Altruistic punishment in humans , author=. Nature , volume=. 2002 , publisher=

  20. [28]

    Nature , volume=

    Evolution of indirect reciprocity , author=. Nature , volume=. 2005 , publisher=

  21. [29]

    Science , volume=

    Cooperation through image scoring in humans , author=. Science , volume=. 2000 , publisher=

  22. [30]

    Nature , volume=

    Reputation helps solve the ‘tragedy of the commons’ , author=. Nature , volume=. 2002 , publisher=

  23. [31]

    Games , volume=

    A review of theoretical studies on indirect reciprocity , author=. Games , volume=. 2020 , publisher=

  24. [32]

    Nature , volume=

    Social norm complexity and past reputations in the evolution of cooperation , author=. Nature , volume=. 2018 , publisher=

  25. [33]

    Proceedings of the national academy of sciences , volume=

    Indirect reciprocity with private, noisy, and incomplete information , author=. Proceedings of the national academy of sciences , volume=. 2018 , publisher=

  26. [34]

    Nature Human Behaviour , volume=

    A unified framework of direct and indirect reciprocity , author=. Nature Human Behaviour , volume=. 2021 , publisher=

  27. [35]

    Nature , volume=

    Development of cooperative relationships through increasing investment , author=. Nature , volume=. 1998 , publisher=

  28. [36]

    Linear reactive strategies , author=

    The continuous Prisoner's Dilemma: I. Linear reactive strategies , author=. Journal of Theoretical Biology , volume=. 1999 , publisher=

  29. [37]

    Nature Communications , volume=

    Quantitative assessment can stabilize indirect reciprocity under imperfect information , author=. Nature Communications , volume=. 2023 , publisher=

  30. [38]

    1985 , publisher=

    Culture and the Evolutionary Process , author=. 1985 , publisher=

  31. [39]

    2011 , publisher=

    Cultural evolution: How Darwinian theory can explain human culture and synthesize the social sciences , author=. 2011 , publisher=

  32. [40]

    Journal of Economic Behavior & Organization , volume=

    Cultural group selection, coevolutionary processes and large-scale cooperation , author=. Journal of Economic Behavior & Organization , volume=. 2004 , publisher=

  33. [41]

    and Filippas, Apostolos and Manning, Benjamin S

    Horton, John J. and Filippas, Apostolos and Manning, Benjamin S. , year =. Large Language Models as Simulated Economic Agents: What Can We Learn from. doi:10.3386/w31122 , url =

  34. [42]

    Proceedings of the National Academy of Sciences , volume=

    A Turing test of whether AI chatbots are behaviorally similar to humans , author=. Proceedings of the National Academy of Sciences , volume=. 2024 , publisher=

  35. [43]

    Proceedings of the International AAAI Conference on Web and Social Media , volume=

    Nicer than humans: how do large language models behave in the prisoner's dilemma? , author=. Proceedings of the International AAAI Conference on Web and Social Media , volume=

  36. [44]

    Scientific Reports , volume=

    Strategic behavior of large language models and the role of game structure versus contextual framing , author=. Scientific Reports , volume=. 2024 , publisher=

  37. [45]

    Political Analysis , volume=

    Out of one, many: Using language models to simulate human samples , author=. Political Analysis , volume=. 2023 , publisher=

  38. [46]

    arXiv preprint arXiv:2403.08882 , year=

    Cultural evolution in populations of Large Language Models , author=. arXiv preprint arXiv:2403.08882 , year=

  39. [47]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Foundations of cooperative AI , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.