Pith. sign in

REVIEW 4 major objections 6 minor 2 cited by

Trust in AI is best understood not as a one-shot adoption choice but as a dynamic decision to monitor less over repeated interactions, and the co-evolution of user and developer strategies converges to just three long-run regimes, only one

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 05:39 UTC pith:KHTZQ5F2

load-bearing objection A genuinely useful formal extension of trust-as-monitoring to an asymmetric AI-governance game, but the abstract overclaims the role of user monitoring and the punishment assumption silently carries the three-regime result. the 4 major comments →

arxiv 2603.24742 v2 pith:KHTZQ5F2 submitted 2026-03-25 cs.AI cs.LGcs.MAnlin.AO

Trust or Check? Understanding the (Evolutionary) Dynamics of User Trust in AI Systems

classification cs.AI cs.LGcs.MAnlin.AO MSC 91A2291A80
keywords AI governancetrust as reduced monitoringevolutionary game theoryreplicator dynamicsreinforcement learninguser trustAI safetymonitoring cost
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes that when user trust is modeled as reduced monitoring in a repeated interaction with AI developers, the co-evolutionary dynamics produce three robust long-run regimes: no adoption alongside unsafe development, unsafe but widely adopted systems, and safe systems that are widely adopted. The only desirable regime—safe and widely adopted—arises exactly when institutional penalties for unsafe behaviour exceed the extra cost of producing safe AI and users can still afford to monitor at least occasionally. The result holds across infinite-population replicator dynamics, stochastic finite-population imitation, and Q-learning agents, making it robust to the choice of learning model. This matters because it formally supports governance proposals centered on transparency, low-cost monitoring, and meaningful sanctions, and it shows that neither regulation alone nor blind user trust is sufficient to prevent drift toward unsafe or low-adoption outcomes.

Core claim

The central claim is that when trust is operationalized as reduced monitoring in a repeated asymmetric game between users and AI developers, the co-evolutionary dynamics converge to exactly three stable long-run regimes in the infinite-population replicator analysis: p4 (AllN, D), where users never adopt and developers produce unsafe AI, stable iff the risk factor μ is negative; p5 (AllA, D), where users always adopt unsafe AI, stable iff μ is positive and institutional punishment v is less than the developer's safety cost c; and p9 (AllA, C), where users always adopt and developers produce safe AI, stable iff v exceeds c. The desirable regime p9 therefore requires institutional punishment t

What carries the argument

The central object is a repeated two-player game between a user and an AI developer, in which a user chooses one of five strategies—AllA (always adopt), AllN (never adopt), TFT (adopt but always monitor, conditioning on the previous outcome), TUA (trust after observing a streak of cooperation, then monitor only with small probability), and DtG (distrust after a streak of defection, then monitor only with small probability)—and the developer chooses C (produce safe/compliant AI, paying cost c) or D (produce unsafe/non-compliant AI, risking institutional punishment v). Trust is operationalized as reduced monitoring: it is the option to stop checking after sufficient evidence of good behavior.

Load-bearing premise

The entire threshold structure rests on treating institutional punishment v as a deterministic per-round payoff loss that a defecting developer pays whenever the user adopts, with no dependence on whether the user actually detects or reports the violation; if punishment were probabilistic or detection-dependent, the stability conditions would need to be re-derived.

What would settle it

A concrete observation that would settle the central claim: find a market or laboratory setting where monitoring is cheap (ε small), penalties exceed safety costs (v > c), and yet unsafe development with wide adoption persists as a stable long-run outcome; or, conversely, a setting where this condition fails and safe, widely adopted AI nonetheless remains stable. Either would contradict the replicator predictions that p9 is the only stable safe regime when v > c and that p5 is stable when μ > 0 and v < c.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Governance should aim to lower the real cost of checking AI systems (ε), since affordable monitoring is a necessary condition for the safe-and-adopted regime to be reachable.
  • Institutional penalties for unsafe AI must exceed the extra cost of producing safe AI (v > c); otherwise the system converges to unsafe-but-adopted or no-adoption outcomes.
  • Blind user trust is dangerous: when unsafe AI still appears beneficial (μ > 0) and punishment is weak (v < c), the system stably settles into wide adoption of unsafe systems.
  • Regulation alone cannot deliver trustworthy AI; user vigilance and meaningful sanctions must work together to prevent evolutionary drift toward unsafe development.
  • The three-regime outcome and the stability thresholds are robust across replicator dynamics, finite-population stochastic imitation, and Q-learning agents.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the model extends to real markets, it yields a testable prediction: policy interventions that reduce verification costs (transparency, standardized audits, accessible evaluation reports) should shift empirical adoption and incident data toward the safe-adopted regime, whereas fines below compliance cost should not.
  • The framework could naturally be extended by making the punishment parameter v endogenous—for example, by adding regulators or auditors as strategic players; the current design keeps them implicit, so the thresholds v > c and v < c would likely become more complex but remain the core ordering principle.
  • A subtle consequence of the risk-dominance analysis is that longer interaction horizons (larger r) amplify the importance of safety costs relative to penalties, suggesting that governance may need to be stricter for long-lived AI products than for one-off deployments.
  • The stability conditions suggest an empirical calibration program: measure ε, c, and v in a given AI market, compute the inequalities, and predict which regime should dominate; laboratory or field data contradicting the prediction would directly test the model.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper studies the co-evolution of user trust and AI developer behaviour in a repeated, asymmetric game, operationalising trust as reduced monitoring. Users choose among unconditional adoption (AllA), non-adoption (AllN), persistent monitoring (TFT), and two threshold-based trust strategies (TUA, DtG); developers choose safe (C) or unsafe (D) development. Payoffs include user benefits, monitoring costs, risk of unsafe AI, safety cost, and an institutional punishment v. The authors analyse the model with finite-population stochastic dynamics, infinite-population replicator dynamics, and Q-learning simulations, and report three robust long-run regimes: no adoption with unsafe development, unsafe but widely adopted systems, and safe systems that are widely adopted. They conclude that the desirable safe regime requires penalties for unsafe behaviour to exceed the extra cost of safety and users to be able to monitor at least occasionally.

Significance. If the conclusions hold, the paper makes a useful conceptual contribution by separating trust (reduced monitoring) from adoption in an asymmetric AI-governance setting, and by showing how monitoring costs and institutional sanctions interact across three dynamical frameworks. The algebraic stability analysis of the corner equilibria in Lemma IV.2 is transparent and parameter-free in the sense that thresholds are derived from the stated payoffs rather than fitted; the finite-population Markov-chain analysis and the Q-learning experiments provide complementary perspectives. However, the central policy claim depends on a modelling choice about institutional punishment that is not currently consistent with the paper's own verbal description, and the 'robust across approaches' claim is stronger than the presented evidence.

major comments (4)
  1. [Section II.A, Table III, Eq. (13), Lemma IV.2] The institutional punishment v is applied as a deterministic per-adoption fine in every D-column entry, including against AllA users who never monitor. This contradicts the text, which says punishment occurs 'when unsafe behaviour is detected' and that users affect creators via 'how often unsafe behaviour is detected (via monitoring and punishment)'. Consequently, the stability condition for p9 (AllA, C) in Lemma IV.2 is simply v>c, with no dependence on monitoring cost epsilon or on user monitoring. The abstract's clause 'and users can still afford to monitor at least occasionally' is therefore not derived in the infinite-population analysis. If punishment were conditional on a monitoring user detecting a violation, the AllA-D payoff would become bc rather than bc-v, and the eigenvalue condition for p9 would no longer reduce to v>c; the three-regime classification and the policy conclus
  2. [Section IV.B, Figure 4, Summary paragraph] The paper claims 'three robust long-run regimes', but the infinite-population analysis establishes only local stability of three corner equilibria (p4, p5, p9). Figure 4 shows, for epsilon=0.5, sustained periodic behaviour between cooperation and defection rather than convergence to any of these pure profiles; the finite-population stationary distributions in Figures 2 and 3 also include substantial mixtures of TFT/TUA/DtG. The statement in the Summary that 'the long-run outcomes are still dominated by these simple pure strategy profiles' is not consistent with these numerical observations. The authors should either restrict the claim to local stability of parameter regimes, or characterise the global attractors and basins of attraction before using the phrase 'robust long-run regimes'.
  3. [Section V.A-B, Eq. (23), Figure 5] The Q-learning analysis is not sufficiently specified to support the 'across these approaches' claim. After stating that 'there is no state transition', the algorithm reduces to a stateless bandit (Eq. 23), but it is not explained how history-dependent strategies such as TFT, TUA, and DtG are represented as actions in this setting. Figure 5 shows only four epsilon values, one fixed parameter set, and averages over 10 runs without error bars or statistical quantification. Moreover, the RL trajectories reported do not themselves exhibit the three claimed regimes; they show a gradual shift from cooperation to defection and AllN as monitoring cost increases. The RL evidence should be treated as illustrative, or the section should be expanded with the missing model specification, sweeps, and repeated-run statistics.
  4. [Lemma IV.2(1)] The proof of Lemma IV.2(1) states that the set p_T cannot be stable because the last eigenvalue is 'bu mu (r-1)/r > 0'. This is not positive for all parameter values: when mu<0, the parameter regime used in all the paper's simulations (e.g., mu=-0.2), this eigenvalue is negative. The argument that these degenerate equilibria cannot be stable is therefore not established for the presented case; the stability of this set requires further analysis or an assessment of the centre-manifold dynamics.
minor comments (6)
  1. [Eq. (22f)] The initial condition is written as (x0, y0, z0, w0, z0); the fifth coordinate should presumably be alpha0 (the initial creator frequency), not z0.
  2. [Table IV] Several risk-dominant conditions are unclear or contain undefined symbols: rows 4 and 5 use 'b' without definition, and rows 8 and 11 include the creator's safety cost c in a transition condition for users. No derivation is provided for these table entries, and at least some appear to have been carried over from a different model. This table should be corrected or removed.
  3. [Section II.B.1] The definition of 'level of adoption' reads 'the frequency when users adopt (i.e. playing T with the creator)'. This is confusing, because the user strategies are AllA, AllN, TFT, TUA, and DtG, not T; the phrase 'playing T' appears to be a leftover from a previous model.
  4. [Section II.A / cross-references] The text refers to 'Tables II A' when describing the payoff matrix; the payoff matrix is Table III. Also, Table II is titled 'Parameters', not 'payoff matrix'.
  5. [Table IV, row 8] The strategy 'AllT' appears instead of 'AllA' (or another defined strategy); this should be corrected for consistency with Table I and the rest of the paper.
  6. [Section V.B, Figure 5] Please state the random seed policy and provide confidence intervals or standard deviations for the 10-run averages; a single selected trajectory set is not sufficient for the text's 'robust' language.

Circularity Check

0 steps flagged

No significant circularity: the three-regime stability conditions are algebraic consequences of the stated payoff matrix; prior-work citations supply assumptions, not imported conclusions.

full rationale

The derivation chain is self-contained at the level that matters for the central claim. The three-regime classification in Section IV is obtained by solving the replicator equilibria (Eqs. 22) and evaluating Jacobian eigenvalues (Table VI): p4 is stable iff μ<0, p5 iff μ>0 and v<c, and p9 iff v>c (Lemma IV.2). These conditions are direct algebraic consequences of the payoff matrix in Table III and of Eq. (13); no parameter is fitted, and no target conclusion is assumed. The trust-as-monitoring premise is imported from the authors' prior work (refs. [13,16]), but it is an explicitly stated modeling assumption, not a conclusion that is then "predicted"; the derived results are conditional on that assumption rather than equivalent to it. The abstract's additional clause that the desirable regime requires users "can still afford to monitor at least occasionally" is not supported by Lemma IV.2, whose p9 stability condition is independent of the monitoring cost ϵ; that is a support/overclaim concern, not a circularity. No uniqueness theorem, hidden ansatz, or renamed empirical pattern is used; self-citations are present but not load-bearing for the algebraic stability analysis.

Axiom & Free-Parameter Ledger

11 free parameters · 7 axioms · 0 invented entities

No new entities are invented; TUA and DtG are strategies, not entities. The model uses many free parameters chosen for illustration rather than fitted to data. The central conclusions are inequalities over these parameters. The most fragile load-bearing axiom is the deterministic punishment structure, which directly drives the v > c condition for the desirable regime.

free parameters (11)
  • monitoring cost ε = 0.0, 0.5, 1.0, 1.5, 2.0 (varied)
    Key bifurcation parameter in the model; swept across simulations to show decline of trust-based strategies.
  • institutional punishment v = 0.1, 0.5, 1.0 (varied)
    Swept in finite-population analysis; the central condition v > c determines the desirable regime.
  • safety cost c = 0.5
    Cost of compliant development; fixed in all numerical figures.
  • risk factor μ = -0.2 (main), 0.2 (appendix)
    Controls harm from unsafe AI; its sign determines stability of p4 versus p5.
  • benefits b_u, b_c = 4, 4
    Symmetric user and creator benefits, chosen for numerical illustration.
  • rounds r = 10
    Length of the repeated interaction; affects payoff averaging and punishment division.
  • thresholds θ_T, θ_D = 3, 3
    Consecutive rounds required before switching to trust or distrust states.
  • checking probabilities p_T, p_D = 0.25, 0.25
    Probability of monitoring in trust/distrust phases of TUA and DtG.
  • selection strength β = 0.1
    Inverse temperature in the Fermi imitation process; set for finite-population simulations.
  • population sizes Z_u, Z_c = 100, 100
    Sizes of the finite user and creator populations.
  • Q-learning rates α, ε_L = 0.05, 0.05
    Learning rate and exploration rate in the reinforcement learning simulations.
axioms (7)
  • domain assumption Trust is operationalized as reduced monitoring (from Perret et al. [13] and Han et al. [16]).
    Adopted as the definition of trust; all user strategies are expressed through adoption and monitoring choices.
  • domain assumption The payoff matrix in Table III correctly captures user and creator incentives, including deterministic punishment v on adoption.
    The entire equilibrium analysis is derived from these payoffs; no empirical validation is provided.
  • domain assumption Populations are well-mixed and evolve by payoff-biased social learning (Fermi process or replicator dynamics).
    Standard EGT setup; excludes network structure, homophily, and strategic forward-looking reasoning.
  • standard math The small-mutation-rate limit justifies the Markov-chain stationary distribution over monomorphic states.
    Standard approximation (Imhof et al. [33], Nowak et al. [34]) used without checking its quantitative validity for the chosen population size of 100.
  • domain assumption Regulators are implicit; punishment v applies deterministically to defecting creators conditional on user adoption.
    Stated in Section II.A; there is no regulator strategy, detection probability, or monitoring-dependent reporting.
  • domain assumption Q-learning agents update in a stateless bandit (Eq. 23), treating the repeated game as a single-shot expected-payoff game.
    This reduction is not justified in the paper; actual repeated-game state histories are ignored.
  • standard math Linear stability analysis via Jacobian eigenvalues suffices to classify equilibria.
    Standard for replicator dynamics; global basins are only studied numerically.

pith-pipeline@v1.3.0-alltime-deepseek · 21014 in / 16686 out tokens · 170992 ms · 2026-08-04T05:39:48.807406+00:00 · methodology

0 comments
read the original abstract

As the capabilities and adoption of Artificial Intelligence (AI) systems grow, trust in these AI systems is an increasingly urgent concern. Much research has focused on models of AI governance and has primarily examined incentives for safe development and effective regulation. Hence they typically represented users trust as a one-shot adoption choice rather than as a dynamic, evolving process shaped by repeated interactions. We instead model trust as the dynamic choice of reduced monitoring in a repeated, asymmetric interaction between users and AI developers, where checking developers' behaviour is costly. Using evolutionary game theory, we study how users' strategies of trust and developers' strategies of providing safe (compliant) or unsafe (non-compliant) AI co-evolve under different levels of monitoring cost and institutional regimes. We conduct the analysis on both imitation-based and learning-based perspectives, with the stochastic finite-population dynamics, the infinite-population replicator analysis and the reinforcement learning analysis. We find three robust long-run regimes: no adoption by users while developers provide unsafe AI, unsafe but widely adopted systems, and safe systems that are widely adopted. Only the last is desirable, and it arises when penalties for unsafe behaviour exceed the extra cost of safety and users can still afford to monitor at least occasionally. Our results formally support governance proposals that emphasise transparency, low-cost monitoring, and meaningful sanctions, and they show that neither regulation alone nor blind user trust is sufficient to prevent the drift towards unsafe or low-adoption outcomes.

Figures

Figures reproduced from arXiv: 2603.24742 by Adeela Bashir, Alessandro Di Stefano, Andrew Powell, Chaimaa Tarzi, Chin-wing Leung, Dhanushka Dissanayake, Elias Fernandez Domingos, Fernando P. Santos, Grace Ibukunoluwa Ufeoshi, Manh Hong Duong, Manuel Chica Serrano, Marcus Krellner, Martin Smit, Nataliya Balabanova, Ndidi Bianca Ogbo, Nikita Huber-Kralj, Paolo Bova, Paolo Turrini, Simon T. Powers, Stefan Sarkadi, The Anh Han, Victor A. Vargas-Perez, Zhao Song, Zia Ush Shamszaman.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4: Numerical modelling of user (top row) and creator (bottom row) cooperation rates for [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5: Percentage of users (top row) adopting different strategies and creator (bottom row) cooperation rates [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Strategic commitments shape collective cybersecurity under AI inequality

    cs.AI 2026-05 unverdicted novelty 5.0

    Subsidized commitment by a small group of defenders in an evolutionary game model significantly increases strong defense adoption, suppresses attacks, and improves system resilience under AI access inequality.

  2. Strategic commitments shape collective cybersecurity under AI inequality

    cs.AI 2026-05 unverdicted novelty 4.0

    Targeted subsidies for committed defenders in an evolutionary game model of AI-unequal cybersecurity significantly increase strong defense adoption, suppress attacks, and enhance overall resilience.

Reference graph

Works this paper leans on

55 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Payoff calculation.We consider two different well-mixed populations of Users and Creators of sizes, respectivelyN u andN c

    Stochastic dynamics for finite populations a. Payoff calculation.We consider two different well-mixed populations of Users and Creators of sizes, respectivelyN u andN c. Letx,y,z,w, and 1−x−y−z−wbe respectively the fraction of users that adopting strategyAllA,AllN,T F T,T U A, andDtG, and letαand 1−αbe respectively the fraction of creators adopting strate...

  2. [2]

    Additionally, we define the frequency of cooperation as P j=C,i∈{AllA,AllN,T F T,T U A,DtG}Λij, and the level of adoption as the frequency when users adopt (i.e

    representing the interaction of two players at a time [36, 37]. Additionally, we define the frequency of cooperation as P j=C,i∈{AllA,AllN,T F T,T U A,DtG}Λij, and the level of adoption as the frequency when users adopt (i.e. playing T with the creator). TABLE IV: Risk-dominant conditions for creator and user strategies. # Risk-dominant transition Conditi...

  3. [3]

    To describe the dynamics, we consider a set ofmdifferent populations (mis some positive integer), which are infinitely large and well-mixed

    Population dynamics for infinite populations: The multi-population replicator dynamics In this section, we recall the framework of the replicator dynamics for multi-populations [38–40]. To describe the dynamics, we consider a set ofmdifferent populations (mis some positive integer), which are infinitely large and well-mixed. Each populationi,i= 1, . . . m...

  4. [4]

    the setp T consists of degenerate equilibria that cannot be stable, since the last eigenvalue is buµ(r−1) r >0

  5. [5]

    the setp D consists of degenerate equilibria that cannot be stable, since the second eigenvalue is equal to ϵ >0

  6. [6]

    This equilibrium corresponds to a situation where creator plays unsafe and user does not adopt

    the pointp 4 is stable if and only ifµ <0. This equilibrium corresponds to a situation where creator plays unsafe and user does not adopt. 12 TABLE V:Coordinates of the equilibrium points Eq. sets x y z w α p1 0 0 1−w ∈[0,1] 0 p2 0 0 ∈[0,1] 0 1 p3 0 0 1 0 0 p4 0 1 0 0 0 p5 1 0 0 0 0 p6 0 0 0 1 0 p7 0 0 0 0 0 p8 0 1 0 0 1 p9 1 0 0 0 1 p10 0 0 0 1 1 p11 0 0...

  7. [7]

    This equilibrium corresponds to a situation where the user adopts even when the creator plays unsafe

    the pointp 5 is stable if and only ifµ >0andv−c <0. This equilibrium corresponds to a situation where the user adopts even when the creator plays unsafe

  8. [8]

    AI Governance Modelling

    the pointp 9 is stable if and only ifc−v <0. This equilibrium corresponds to a situation where the user always adopts when the creator plays safe. B. Numerical analysis Figure 4 depicts the numerical plots for the ratios of AllA, AllN, TFT, TUA and DtG strategists among users, as well as the percentage of cooperators among creators for varying values of t...

  9. [9]

    We thus assume that users can adopt one of the following strategies: AllA, AllN or TFT

    Evolutionary dynamics in the absence of trust-based strategies In this section, we analyse the model without the trust-based strategies TUA and DtG. We thus assume that users can adopt one of the following strategies: AllA, AllN or TFT. Adopting the latter still requires them to pay a costϵfor checking the action of creators in the previous round. We now ...

  10. [10]

    Trust AI Regulation? Discerning Users are Vital to Build Trust and Effective AI Regulation,

    Z. Alalawi, P. Bova, T. Cimpeanu, A. Di Stefano, M. Hong Duong, E. F. Domingos, T. A. Han, M. Krellner, N. B. Ogbo, S. T. Powers, and F. Zimmaro, “Trust AI Regulation? Discerning Users are Vital to Build Trust and Effective AI Regulation,”Applied Mathematics and Computation, vol. 508, p. 129627, 2026

  11. [11]

    Can Media Act as a Soft Regulator of Safe AI Development? a Game Theoretical Analysis,

    H. C. da Fonseca, A. Fernandes, Z. Song, T. Cimpeanu, N. Balabanova, A. Bashiret al., “Can Media Act as a Soft Regulator of Safe AI Development? a Game Theoretical Analysis,”ALIFE 2025: Ciphers of Life. Proceedings of the Artificial Life Conference 2025, p. 90, 10 2025

  12. [12]

    Media and responsible AI governance: a game-theoretic and LLM analysis,

    N. Balabanova, A. Bashir, P. Bova, A. Buscemi, T. Cimpeanu, H. C. d. Fonseca, A. D. Stefano, M. H. Duong, E. F. Domingos, A. Fernandes, T. A. Han, M. Krellner, N. B. Ogbo, S. T. Powers, D. Proverbio, F. P. Santos, Z. U. Shamszaman, and Z. Song, “Media and responsible AI governance: a game-theoretic and LLM analysis,” Mar. 2025, arXiv:2503.09858 [cs]. [Onl...

  13. [13]

    Do LLMs Trust AI Regulation? Emerging Behaviour of Game-theoretic LLM Agents,

    A. Buscemi, D. Proverbio, P. Bova, N. Balabanova, A. Bashir, T. Cimpeanu, H. C. da Fonseca, M. H. Duong, E. F. Domingos, A. M. Fernandeset al., “Do LLMs Trust AI Regulation? Emerging Behaviour of Game-theoretic LLM Agents,”arXiv preprint arXiv:2504.08640, 2025

  14. [14]

    Trustworthy artificial intelligence and the european union ai act: On the conflation of trustworthiness and acceptability of risk,

    J. Laux, S. Wachter, and B. Mittelstadt, “Trustworthy artificial intelligence and the european union ai act: On the conflation of trustworthiness and acceptability of risk,”Regulation & Governance, vol. 18, no. 1, pp. 3–32, 2024

  15. [15]

    The neural computation of trust and reputation,

    E. Fouragnanet al., “The neural computation of trust and reputation,” 2013

  16. [16]

    Cooperation through indirect reciprocity in child-robot interac- tions,

    I. Neto, A. S. Pires, F. Correia, and F. P. Santos, “Cooperation through indirect reciprocity in child-robot interac- tions,”arXiv preprint arXiv:2512.20621, 2025

  17. [17]

    The trust paradox: a survey of economic inquiries into the nature of trust and trustworthiness,

    H. S. James Jr., “The trust paradox: a survey of economic inquiries into the nature of trust and trustworthiness,” Journal of Economic Behavior & Organization, vol. 47, no. 3, pp. 291–307, Mar. 2002. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167268101002141

  18. [18]

    Separating Trust from Cooperation in a Dynamic Relationship: Prisoner’s Dilemma with Variable Dependence,

    T. Yamagishi, S. Kanazawa, R. Mashima, and S. Terai, “Separating Trust from Cooperation in a Dynamic Relationship: Prisoner’s Dilemma with Variable Dependence,”Rationality and Society, vol. 17, no. 3, pp. 275–308, Aug. 2005, publisher: SAGE Publications Ltd. [Online]. Available: https://doi.org/10.1177/1043463105055463

  19. [19]

    How do we even trust? A critique of economic approaches on trust formation,

    M. Banu, “How do we even trust? A critique of economic approaches on trust formation,”International Journal of Economic Behavior (IJEB), vol. 14, no. 1, pp. 23–38, May 2024. [Online]. Available: https://journals.uniurb.it/index.php/ijmeb/article/view/4541

  20. [20]

    Trust, reciprocity, and social history,

    J. Berg, J. Dickhaut, and K. McCabe, “Trust, reciprocity, and social history,”Games and Economic Behavior, vol. 10, pp. 122–142, 1995

  21. [21]

    The social structure of trust,

    V. Buskens, “The social structure of trust,”Social Networks, vol. 20, no. 3, pp. 265–289, Jul. 1998. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0378873398000057

  22. [22]

    Disentangling trust from cooperation: Evolution of trust as reduced monitoring in social dilemmas,

    C. Perret, T. A. Han, E. F. Domingos, T. Cimpeanu, and S. T. Powers, “Disentangling trust from cooperation: Evolution of trust as reduced monitoring in social dilemmas,”Chaos, Solitons & Fractals, vol. 208, p. 118130, 2026

  23. [23]

    Luhmann,Trust and Power

    N. Luhmann,Trust and Power. Chichester: John Wiley & Sons, 1979

  24. [24]

    In AI we trust incrementally: a multi-layer model of trust to analyze human- artificial intelligence interactions,

    A. Ferrario, M. Loi, and E. Vigan` o, “In AI we trust incrementally: a multi-layer model of trust to analyze human- artificial intelligence interactions,”Philosophy & Technology, vol. 33, no. 3, pp. 523–539, Sep. 2020

  25. [25]

    When to (or not to) trust intelligent machines: Insights from an evolutionary game theory analysis of trust in repeated games,

    T. A. Han, C. Perret, and S. T. Powers, “When to (or not to) trust intelligent machines: Insights from an evolutionary game theory analysis of trust in repeated games,”Cognitive Systems Research, vol. 68, pp. 111–124, 2021

  26. [26]

    Trust does not need to be human: it is possible to trust medical AI,

    A. Ferrario, M. Loi, and E. Vigan` o, “Trust does not need to be human: it is possible to trust medical AI,”Journal of Medical Ethics, vol. 47, no. 6, pp. 437–438, Jun. 2021

  27. [27]

    Being pragmatic about reliance and trust in artificial intelligence,

    A. Ferrario, “Being pragmatic about reliance and trust in artificial intelligence,”Minds and Machines, vol. 36, no. 1, p. 5, Dec. 2025

  28. [28]

    A game-theoretic model of trust in human–robot teaming: Guiding human observation strategy for monitoring robot behavior,

    Z. Zahedi, S. Sengupta, and S. Kambhampati, “A game-theoretic model of trust in human–robot teaming: Guiding human observation strategy for monitoring robot behavior,”IEEE Transactions on Human-Machine Systems, vol. 55, no. 1, pp. 37–47, Feb. 2025

  29. [29]

    How much do you trust me? A logico-mathematical analysis of the concept of the intensity of trust,

    M. Loi, A. Ferrario, and E. Vigan` o, “How much do you trust me? A logico-mathematical analysis of the concept of the intensity of trust,”Synthese, vol. 201, no. 6, p. 186, May 2023

  30. [30]

    Multiagent reinforcement learning in the iterated prisoner’s dilemma,

    T. W. Sandholm and R. H. Crites, “Multiagent reinforcement learning in the iterated prisoner’s dilemma,”Biosys- tems, vol. 37, no. 1-2, pp. 147–166, 1996

  31. [31]

    Open problems in cooperative ai,

    A. Dafoe, E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel, “Open problems in cooperative ai,”arXiv preprint arXiv:2012.08630, 2020

  32. [32]

    Investigating the impact of direct punishment on the emergence of cooperation in multi-agent reinforcement learning systems,

    N. Dasgupta and M. Musolesi, “Investigating the impact of direct punishment on the emergence of cooperation in multi-agent reinforcement learning systems,”Autonomous Agents and Multi-Agent Systems, vol. 39, no. 1, pp. 1–37, 2025

  33. [33]

    Multi-agent reinforcement learning in se- quential social dilemmas,

    J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel, “Multi-agent reinforcement learning in se- quential social dilemmas,” inProceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems (AAMAS’17), 2017

  34. [34]

    Learning fair cooperation in mixed-motive games with indirect reciprocity,

    M. Smit and F. P. Santos, “Learning fair cooperation in mixed-motive games with indirect reciprocity,” inProceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI’24), 2024

  35. [35]

    Collective cooperative intelligence,

    W. Barfuss, J. Flack, C. S. Gokhale, L. Hammond, C. Hilbe, E. Hughes, J. Z. Leibo, T. Lenaerts, N. Leonard, S. Levinet al., “Collective cooperative intelligence,”Proceedings of the National Academy of Sciences, vol. 122, no. 25, p. e2319948121, 2025

  36. [36]

    A multi- agent reinforcement learning model of common-pool resource appropriation,

    J. P´ erolat, J. Z. Leibo, V. F. Zambaldi, C. Beattie, K. Tuyls, and T. Graepel, “A multi- agent reinforcement learning model of common-pool resource appropriation,” inAdvances in Neural Information Processing Systems, I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 3643–3652. [Online]....

  37. [37]

    Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents,

    E. Tennant, S. Hailes, and M. Musolesi, “Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents,” inProceedings of the 7th AAAI/ACM Conference on AI, Ethics, and Society (AIES 2024), 2024

  38. [38]

    Reinforcement learning in evolutionary game theory: A brief review of recent develop- ments,

    K. Xie and A. Szolnoki, “Reinforcement learning in evolutionary game theory: A brief review of recent develop- ments,”Applied Mathematics and Computation, vol. 510, p. 129685, 2026

  39. [39]

    Decoding trust: A reinforcement learning perspective,

    G. Zheng, J. Zhang, J. Zhang, W. Cai, and L. Chen, “Decoding trust: A reinforcement learning perspective,”New Journal of Physics, vol. 26, no. 5, p. 053041, 2024

  40. [40]

    D. C. North,Institutions, institutional change and economic performance. Cambridge university press, 1990

  41. [41]

    Stochastic Dynamics of Invasion and Fixation,

    A. Traulsen, M. A. Nowak, and J. M. Pacheco, “Stochastic Dynamics of Invasion and Fixation,”Physical Review E, vol. 74, p. 11909, 2006

  42. [42]

    Evolutionary Cycles of Cooperation and Defection,

    L. A. Imhof, D. Fudenberg, and M. A. Nowak, “Evolutionary Cycles of Cooperation and Defection,”Proc. Natl. Acad. Sci. U.S.A., vol. 102, pp. 10 797–10 800, 2005

  43. [43]

    Emergence of Cooperation and Evolutionary Stability in Finite Populations,

    M. A. Nowak, A. Sasaki, C. Taylor, and D. Fudenberg, “Emergence of Cooperation and Evolutionary Stability in Finite Populations,”Nature, vol. 428, pp. 646–650, 2004

  44. [44]

    EGTtools: Evolutionary Game Dynamics in Python,

    E. F. Domingos, F. C. Santos, and T. Lenaerts, “EGTtools: Evolutionary Game Dynamics in Python,”Iscience, vol. 26, no. 4, 2023

  45. [45]

    Paradigm Shifts and the Interplay Between State, Business and Civil Sectors,

    S. Encarna¸ c˜ ao, F. P. Santos, F. C. Santos, V. Blass, J. M. Pacheco, and J. Portugali, “Paradigm Shifts and the Interplay Between State, Business and Civil Sectors,”Royal Society open science, vol. 3, no. 12, p. 160753, 2016

  46. [46]

    Pathways to Good Healthcare Services and Patient Satisfaction: An Evolutionary Game Theoretical Approach,

    Z. Alalawi, T. A. Han, Y. Zeng, and A. Elragig, “Pathways to Good Healthcare Services and Patient Satisfaction: An Evolutionary Game Theoretical Approach,” inArtificial Life Conference Proceedings. MIT Press, 2019, pp. 135–142

  47. [47]

    Evolutionarily Stable Strategies with Two Types of Player,

    P. D. Taylor, “Evolutionarily Stable Strategies with Two Types of Player,”Journal of applied probability, vol. 16, no. 1, pp. 76–83, 1979

  48. [48]

    The Stabilization of Equilibria in Evolutionary Game Dynamics Through Mutation: Mutation Limits in Evolutionary Games,

    J. Bauer, M. Broom, and E. Alonso, “The Stabilization of Equilibria in Evolutionary Game Dynamics Through Mutation: Mutation Limits in Evolutionary Games,”Proceedings of the Royal Society A, vol. 475, no. 2231, p. 20190355, 2019

  49. [49]

    Co-evolutionary dynamics of attack and defence in cybersecurity,

    A. Bashir, Z. U. Shamszaman, Z. Song, and T. A. Han, “Co-evolutionary dynamics of attack and defence in cybersecurity,”Knowledge-Based Systems, p. 115750, 2026

  50. [50]

    Emergence of Cooperation and Evolutionary Stability in Finite Populations,

    M. A. Nowak, A. Sasaki, C. Taylor, and D. Fudenberg, “Emergence of Cooperation and Evolutionary Stability in Finite Populations,”Nature, vol. 428, no. 6983, pp. 646–650, 2004

  51. [51]

    Evolution of Fairness in the One-shot Anonymous Ultimatum Game,

    D. G. Rand, C. E. Tarnita, H. Ohtsuki, and M. A. Nowak, “Evolution of Fairness in the One-shot Anonymous Ultimatum Game,”Proceedings of the National Academy of Sciences, vol. 110, no. 7, pp. 2581–2586, 2013

  52. [52]

    Generosity Motivated by Acceptance- Evolutionary Analysis of an Anticipation Game,

    I. Zisis, S. Di Guida, T. A. Han, G. Kirchsteiger, and T. Lenaerts, “Generosity Motivated by Acceptance- Evolutionary Analysis of an Anticipation Game,”Scientific reports, vol. 5, no. 1, p. 18076, 2015

  53. [53]

    Hofbauer and K

    J. Hofbauer and K. Sigmund,Evolutionary Games and Population Dynamics. Cambridge university press, 1998

  54. [54]

    Q-learning,

    C. J. Watkins and P. Dayan, “Q-learning,”Machine Learning, vol. 8, pp. 279–292, 1992

  55. [55]

    Social physics in the age of artificial intelligence,

    T. A. Han, J. Z. Leibo, T. Lenaerts, I. Rahwan, F. Santos, M. Perc, and V. Capraro, “Social physics in the age of artificial intelligence,” 2026. [Online]. Available: https://arxiv.org/abs/2603.16900