REVIEW 3 major objections 3 minor 36 references
Incentives to Build Houses, Trade Houses, or Trade House Building Skills in Simulated Worlds under Various Governing Systems or Institutions: Comparing Multi-agent Reinforcement Learning to Generative Agent-based Model
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A semi-libertarian, vote-counting government with an inclusive institution produces the highest build-to-trade ratios in a simulated economy; in an LLM-based world, Full-Utilitarian wins under equality goals and Full-Libertarian under…
desk verdict A genuinely novel MARL-vs-GABM comparison undermined by single-run pooling, but honest and worth a thorough revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the extended AI-Economist, a two-level deep MARL setup in which a central planner sets tax rates and invests tax revenue while mobile agents vote on resource priorities and choose among building houses, trading houses, and trading skill; and the extended Concordia, a generative agent-based model in which LLM-driven agents act through natural language and a game master translates actions, tracks grounded variables, and sets taxes to maximize equality or productivity.
What would settle it
Run multiple independent seeds for each governing system and institution in both frameworks while holding the government's objective fixed, then recompute the build-to-trade and build-to-skill-trade ratios; if the Semi-Libertarian/Utilitarian and Inclusive ordering does not persist across seeds, the central claim is settled.
Extended reading notes
Core claim
The paper extends the AI-Economist MARL framework so six agents gather wood, stone, and iron, build three house types, trade houses, and trade house building skill, with expert and novice agents distinguished by payment multipliers. It also extends Concordia so LLM-driven agents perform the same activities under a game master that tracks inventory, skill, build, vote, and tax. The central discovery is that governance structure changes which activity agents favor: in the MARL world, the Semi-Libertarian/Utilitarian system, where agents vote and the planner follows their Borda ranking, gives the highest ratios of building houses to trading houses and to trading skill, and among its institutions Inclusive gives higher ratios than Arbitrary and Extractive. In the Concordia world, an equality-optimizing game master makes Full-Utilitarian agents build and trade skill more, while a productivity-optimizing game master makes Full-Libertarian agents build, trade houses, and trade skill more. The paper also reports that the three economic activities correlate positively with productivity, equality, and maximin, with one exception: building houses correlates negatively with equality.
Load-bearing premise
The load-bearing premise is statistical: because each configuration was run once and the displayed ratios pool two runs that differ in the government's objective, the ranking of governing systems could flip if the reward difference or the single draw of randomness is responsible for the pattern.
Editorial extensions
If this is right
- Under the extended AI-Economist, the Semi-Libertarian/Utilitarian system, the closest to current democratic systems, yields the highest build-to-trade and build-to-skill-trade ratios among the three governing systems.
- Within that governing system, the Inclusive institution yields higher build-to-trade and build-to-skill-trade ratios than the Arbitrary and Extractive institutions.
- In the extended Concordia, an equality-optimizing game master makes Full-Utilitarian agents build more houses and trade more house building skill, while a productivity-optimizing game master makes Full-Libertarian agents build more houses, trade more houses, and trade more skill.
- In both frameworks, building, trading houses, and trading skill mostly correlate positively with productivity, equality, and maximin, except that building houses correlates negatively with equality.
- Both MARL and GABM agents partially infer the implicit rules of the environment, suggesting both approaches can model similar social phenomena despite their different architectures.
Reading between the lines
- The pooled two-reward-function design leaves open that the governance ranking is an artifact of the government's objective; disentangling the two reward functions in the plots would test this directly.
- If the ranking survives a random-seed sweep, governance-system simulations could become a low-cost screen for institutional designs before field experiments.
- The negative building-versus-equality correlation found in both frameworks suggests a possible trade-off between construction activity and equality that the paper does not develop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends the AI-Economist MARL framework and the Concordia GABM framework so that agents can build houses, trade houses, and trade house building skill, and it introduces three governing systems (Full-Libertarian, Semi-Libertarian/Utilitarian, Full-Utilitarian) and, in the MARL variant, three governing institutions (Inclusive, Arbitrary, Extractive). The headline claims are that in the extended AI-Economist, the Semi-Libertarian/Utilitarian system and the Inclusive institution produce the highest ratios of building to trading houses and of building to trading skill (Figs. 4-5), and that in the extended Concordia, Full-Utilitarian produces more building and skill trading when the game master maximizes equality, while Full-Libertarian produces more of all three activities when it maximizes productivity (Figs. 11-12). The paper also discusses correlations between economic activities and welfare measures and offers a qualitative comparison of MARL and GABM capabilities, while explicitly acknowledging that only one simulation per parameter set was run.
Significance. If the headline results held, the paper would offer an interesting cross-framework comparison of how governance shapes economic incentives in simulated societies, and it would be a useful exploratory contribution with practical value for in-silico social science. The author provides open-source code for both frameworks, which is a genuine strength and enables reproducibility. However, the statistical foundation of the central comparisons is too weak: with a single simulation per configuration, no seed replicates, no error bars, and pooling of runs that differ in the planner's reward function, the ordering of governing systems is not established. The paper itself acknowledges this limitation, but the abstract and results sections still present the rankings as findings, which overstates the evidence. The GABM results are described in deliberately hedged language and are based on visual inspection, further reducing the strength of the second main claim.
major comments (3)
- [§3, Figs. 4-5, Fig. 15, §4 Current Limitations] The central MARL claim—that Semi-Libertarian/Utilitarian and Inclusive produce higher build/trade ratios—is computed by averaging over two runs per condition that differ in the central planner's reward function (Fig. 15 caption, Figs. 4-5). The author explicitly states in §4 that 'the number of simulations for each set of input parameters ... is one,' so these two runs are not replicates but distinct experimental conditions. Because the planner's reward function determines tax policy and agents' build/trade decisions respond to taxes, pooling such runs can create a spurious ordering even if neither reward function individually supports it. This is a load-bearing issue: without separating these conditions or adding proper replicates, the headline ranking is not supported by the presented data.
- [§3, Figs. 4-5, 11-12] No seed replicates, error bars, or significance tests are provided for any condition in either framework. With a single trajectory per parameter set in the MARL experiments, the observed ratios could flip under a different random seed; the paper notes that two-level RL training is 'particularly unstable' (Fig. 16 caption), which makes the absence of multiple runs especially concerning. For the GABM results, claims are made from visual inspection of ten episodes, with the author using hedged language ('it seems') and reporting one 'hallucination' (Fig. 10). The lack of any uncertainty quantification means that even the qualitative ordering of governance systems is unverified.
- [§2.2, §3, Figs. 10-12] The Concordia results constitute a second main claim of the abstract, yet they are reported only qualitatively. Figure 11 is interpreted as showing that Full-Utilitarian leads to 'slightly higher' building and skill trading under equality, and Figure 12 as showing that Full-Libertarian leads to higher activity under productivity, but no quantitative measure, statistical comparison, or sensitivity analysis is given. This is not a presentation issue but a limitation of the evidence for a central claim: the paper's own description ('it seems') acknowledges that the results are not robustly demonstrated.
minor comments (3)
- [§4 Conclusions] The word 'fro' appears in 'fro both MARL and GABM'; a typo for 'for'. Several figure captions also contain grammatical errors, e.g., 'At it is clear form these plots' (Figs. 11-12) and 'refering' (Fig. 7).
- [§2.1] The text refers to 'the four resources' when describing the Semi-Libertarian/Utilitarian planner's investment of tax revenue, but the environment has only three material resources (wood, stone, and iron); this is inconsistent with the rest of the paper and should be corrected to 'three resources.'
- [Abstract and §1] The phrase 'somewhat similar worlds' is appropriate given the differences between the MARL and GABM implementations, but the paper should explicitly state that the two frameworks are not directly comparable in a quantitative sense, since one uses learned policies and the other uses LLM prompting. The current language in the Introduction ('as similar as possible') is still vague; a short paragraph detailing the key differences (e.g., learning dynamics, action spaces, observation grounding) would strengthen the comparison.
Circularity Check
No circularity: the reported build/trade ratios are simulation outputs, not derived identities or fitted predictions.
full rationale
The paper's central claims are measured ratios and activity counts obtained by running the extended AI-Economist and Concordia under different governing systems and institutions. These are empirical results of simulation episodes (e.g., Figs. 4, 5, 11, 12), not conclusions derived from an equation that is equivalent to its own inputs. The governing-system definitions and extensions are attributed to the author's prior work (Dizaji 2023a,b, 2024), but the current build/trade ordering is not inferred from those prior results; it comes from new simulation runs whose code is available. No parameter is fitted to a subset of the data and then renamed as a prediction, no uniqueness theorem is imported to force a choice, and no ansatz is smuggled in via self-citation. The pooling of two runs differing in the planner's reward function (Fig. 15) and the single-run limit are genuine methodological/statistical limitations, and the paper itself acknowledges them in the Current Limitations section, but they are not circularity. The Concordia findings are hedged as 'it seems' and summarized from ten LLM episodes, again observational rather than derivational. No step in the claimed derivation chain reduces to its own inputs by construction.
Assumptions & free parameters
free parameters (3)
- expert/novice payment multiplier means =
not stated in text (in configuration figures)
- minimum payment multiplier threshold for building =
fixed value
- labor cost for trading (zero) vs building (positive) =
zero for trades
assumptions (4)
- domain assumption PPO with shared weights converges to meaningful policies in 5000 steps
- domain assumption LLM-based Concordia agents can track inventories and make coherent economic decisions through natural language
- domain assumption The MARL and GABM environments are comparable enough for a qualitative comparison
- ad hoc to paper Pooling two runs with different planner reward functions yields a valid estimate of the ratios
invented entities (1)
-
house building skill
Cite this review
Pith. "Pith review of Incentives to Build Houses, Trade Houses, or Trade House Building Skills in Simulated Worlds under Various Governing Systems or Institutions: Comparing Multi-agent Reinforcement Learning to Generative Agent-based Model." pith.science (2026). https://pith.science/paper/UB7HJ6QS
@misc{pith2026241117724,
author = {Pith},
title = {Pith review of: Incentives to Build Houses, Trade Houses, or Trade House Building Skills in Simulated Worlds under Various Governing Systems or Institutions: Comparing Multi-agent Reinforcement Learning to Generative Agent-based Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/UB7HJ6QS}},
note = {Machine review of arXiv:2411.17724}
}
read the original abstract
It has been shown that social institutions impact human motivations to produce different behaviours, such as amount of working or specialisation in labor. With advancement in artificial intelligence (AI), specifically large language models (LLMs), now it is possible to perform in-silico simulations to test various hypotheses around this topic. Here, I simulate two somewhat similar worlds using multi-agent reinforcement learning (MARL) framework of the AI-Economist and generative agent-based model (GABM) framework of the Concordia. In the extended versions of the AI-Economist and Concordia, the agents are able to build houses, trade houses, and trade house building skill. Moreover, along the individualistic-collectivists axis, there are a set of three governing systems: Full-Libertarian, Semi-Libertarian/Utilitarian, and Full-Utilitarian. Additionally, in the extended AI-Economist, the Semi-Libertarian/Utilitarian system is further divided to a set of three governing institutions along the discriminative axis: Inclusive, Arbitrary, and Extractive. Building on these, I am able to show that among governing systems and institutions of the extended AI-Economist, under the Semi-Libertarian/Utilitarian and Inclusive government, the ratios of building houses to trading houses and trading house building skill are higher than the rest. Furthermore, I am able to show that in the extended Concordia when the central government care about equality in the society, the Full-Utilitarian system generates agents building more houses and trading more house building skill. In contrast, these economic activities are higher under the Full-Libertarian system when the central government cares about productivity in the society. Overall, the focus of this paper is to compare and contrast two advanced techniques of AI, MARL and GABM, to simulate a similar social phenomena with limitations.
Figures
Figures from the paper (28 more)
Reference graph
Works this paper leans on
-
[1]
The AI-Economist is a two-level deep RL framework for policy design in which agents and a social planner co-adapt. In particular, the AI-Economist uses structured curriculum learning to stabilize the challenging two-level, co-adaptive learning problem. This framework has been validated in the domain of taxation. In two-level spatiotemporal economies, the ...
-
[2]
Stabilizing the training process in two-level RL is difficult. To overcome, the training procedure in the AI-Economist has two important features - curriculum learning and entropy-based regularization. Both of them encourage the agents and the social planner to co-adopt gradually and not stopping exploration too early during training and getting trapped i...
-
[3]
The Gather-Trade-Build economy of the AI-Economist is a two-dimensional spatiotemporal economy with agents who move, gather resources (stone and wood), trade them, and build houses. Each agent has a varied house build-skill which sets how much income an agent receives from building a house. Build-skill is distributed according to a Pareto distribution. As...
-
[4]
The Open-Quadrant environment of the Gather-Trade-Build economy has four regions delineated by impassable water with passageways connecting each quadrant. Quadrants contain different combi- nations of resources: both stone and wood, only stone, only wood, or nothing. Agents can freely access all quadrants, if not blocked by objects or other agents. This s...
-
[5]
The action space of the agents includes four movement actions: up, down, left, and right
The state of the world is represented as an nh × nw × nc tensor, where nh and nw are the size of the world and nc is the number of unique entities that may occupy a cell, and the value of a given element indicates which entity is occupying the associated location. The action space of the agents includes four movement actions: up, down, left, and right. Ag...
-
[6]
Agent’s observations include the state of their own endowment (wood, stone, and coin), their own build-skill level, and a view of the world state tensor within an egocentric spatial window. The experiment use a world of 25 by 25 for 4-agent and 40 by 40 for 10-agent environments, where agent spatial observations have size 11 by 11 and are padded as needed...
-
[7]
Agents can buy and sell resources from one another through a continuous double-auction. Agents can submit asks (the number of coins they are willing to accept) or bids (how much they are willing to pay) in exchange for one unit of wood or stone. The action space of the agents includes 44 actions for trading, representing the combination of 11 price levels...
-
[8]
Agents can choose to spend one unit of wood and one unit of stone to build a house, and this places a house tile at the agent’s current location and earns the agent some number of coins. Agents are restricted from building on source cells as well as locations where a house already exists. The number of coins earned per house is identical to an agent’s bui...
Show all 36 references
-
[9]
Taxation is implemented using income brackets and bracket tax rates
Simulations are run in episodes of 1000 time steps, which is subdivided into 10 tax periods or tax years, each lasting 100 time steps. Taxation is implemented using income brackets and bracket tax rates. All taxation is anonymous: Tax rates and brackets do not depend on the id...
-
[10]
The payable tax for income z is computed as follows: T (z) = BX j=1 τj · ((bj+1 − bj)1[z > bj+1] + (z − bj)1[bj < z≤ bj+1]) (1) where B is the number of brackets, and τj and bj are marginal tax rates and income boundaries of the brackets, respectively
-
[11]
Accordingly, taxes are collected at the end of each tax year by subtracting T (zi) from Ci
An agent’s pretax income zi for a given tax year is defined simply as the change in its coin endowment Ci since the start of the year. Accordingly, taxes are collected at the end of each tax year by subtracting T (zi) from Ci. Taxes are used to redistribute wealth: the total t...
-
[12]
It is found that build-skill is a substantial determinant of behavior; agents’ gather-skill empirically does not affect optimal behaviors in our settings
Agents learn behaviors that maximize their expected total discounted utility for an episode. It is found that build-skill is a substantial determinant of behavior; agents’ gather-skill empirically does not affect optimal behaviors in our settings. All of the experiments use a ...
2018
-
[13]
RL is instantiated at two levels, that is, for two types of actors: training agent behavioral policy models and a taxation policy model for the social planner
RL provides a flexible way to simultaneously optimize and model the behavioral effects of tax policies. RL is instantiated at two levels, that is, for two types of actors: training agent behavioral policy models and a taxation policy model for the social planner. Each actor’s ...
-
[14]
To improve learning efficiency, a single-agent policy networkπ(ai,t|oi,t; θ) is trained whose weights are shared by all agents, that is, θi = θ
In this work, the proximal policy gradients (PPO) is used to train all actors (both agents and planner). To improve learning efficiency, a single-agent policy networkπ(ai,t|oi,t; θ) is trained whose weights are shared by all agents, that is, θi = θ. This network is still able ...
-
[15]
At each time step t, each agent observes the following: its nearby spatial surroundings; its current endowment (stone, wood, and coin); private characteristics, such as its building skill; the state of the markets for trading resources; and a description of the current tax rat...
-
[16]
Rational economic agents train their policy πi to optimize their total discounted utility over time while experiencing tax rates τ set by the planner’s policy πp. The agent training objective is: ∀i : max πi Eτ ∼πp,ai∼πi,a−i∼π−i,s′ ∼P [ HX t=1 γtri,t + ui,0], ri,t = ui,t − ui,...
-
[17]
For an agent population with monetary endowments Ct = (C1,t, ..., CN,t), the equality eq(Ct) is defined as: eq(Ct) = 1− N N − 1 gini(Ct), 0 ≤ eq(Ct) ≤ 1 (6) where the Gini index is defined as: gini(Ct) = PN i=1 PN j=1 |Ci,t − Cj,t| 2N PN i=1 Ci,t , 0 ≤ gini(Ct) ≤ N − 1 N (7)
-
[18]
Hence, the sum of pretax and post-tax incomes is the same
The productivity is defined as the sum of all incomes: prod(Ct) = X i Ci,t (8) The economy is closed: subsidies are always redistributed evenly among agents, and no tax money leaves the system. Hence, the sum of pretax and post-tax incomes is the same. The planner trains its p...
-
[19]
The utilitarian social welfare objective is the family of linear-weighted sums of agent utilities, defined for weights ωi ≥ 0: swft = NX i=1 ωi · ui,t (10) 20 Inverse-income is used as the weights: ωi ∝ 1 Ci , normalized to sum to one. An objective function is defined that opt...
2023
-
[20]
In Concordia both are generative
A generative modelling of social interactions have two parts: the model of the environment and the model of individual behaviour. In Concordia both are generative. Thus in Concordia there are : (1) generative agents and (2) a generative model for the environment, space, or wor...
-
[21]
The game master receives the agent actions and produces event statements, which define the course of events in the simulation as a result of the agent’s generated action
Concordia agents receive observations of the environment as inputs and based on those generate actions. The game master receives the agent actions and produces event statements, which define the course of events in the simulation as a result of the agent’s generated action. Th...
-
[22]
The game master absorbs their intended actions, decides on the outcome of their attempts, and generates event statements
Concordia agents generate their behaviours by explaining what they want to do in natural language. The game master absorbs their intended actions, decides on the outcome of their attempts, and generates event statements. Basically, the game master is performing the following t...
-
[23]
The game master determines the effect of the agents’ actions on these variables, records them, and checks that they are valid
The game master’s most important responsibility is to provide the grounding for particular experimental variables, which are defined for an experiment. The game master determines the effect of the agents’ actions on these variables, records them, and checks that they are valid...
-
[24]
The produced agents behaviours should be consistent with common sense, in accordance by social norms, and individually grounded based on a personal history of past events as well as ongoing understanding of the current situation
-
[25]
It is argued that humans generally act as though they choose their actions by answering three key questions: (1) What kind of situation is this? (2) What kind of person am I? (3) What does a person such as I do in a situation such as this?
-
[26]
The idea is that, if the outputs of LLMs conditioned to model specific human populations, they reflect the beliefs and attitudes of those populations
The premise behind Concordia is that since modern LLMs have been trained on huge amounts of human culture, they are thus capable of giving reasonable answers to the above questions when provided with the historical context of a particular agent. The idea is that, if the output...
-
[27]
Concordia makes this possible by using an associative memory in a modular and flexible fashion to keep the record of agents experience
The next step is to make available a record of an agent’s historical experience to an LLM so it would be able to answer the above mentioned key questions. Concordia makes this possible by using an associative memory in a modular and flexible fashion to keep the record of agent...
-
[28]
The working memory is zi i composed of the states of individual components
Memory is a set of strings m that records everything remembered or currently experienced by the agent. The working memory is zi i composed of the states of individual components. A component i has a state zi, which is statement in natural language. The components update their ...
-
[29]
When creating a generative agent in Concordia, the user creates the components that are relevant for their simulations
The incoming observations are fed into the agents memory to make them available when components update. When creating a generative agent in Concordia, the user creates the components that are relevant for their simulations. They decide on the initial state and the update funct...
-
[30]
The most simple form of fa is a concatenation operator over zt = zi t i
Here fa is a formatting function, which out of the states of components creates the grounding used to sample the action to take. The most simple form of fa is a concatenation operator over zt = zi t i. We do not explicitly condition on the memory m or observation o, since we c...
-
[31]
In Concordia, conditioning is explicitely done on the memory stream m, since a component may make specific queries into the agent’s memory to update its state
Here, f i is a formatting function that turns the memory stream and the current state of the components into the query for the component update. In Concordia, conditioning is explicitely done on the memory stream m, since a component may make specific queries into the agent’s ...
-
[32]
The game master mediates between the state of the world and agents’ actions
The game master takes care of all aspects of the simulated world not directly controlled by the agents. The game master mediates between the state of the world and agents’ actions. The state of the world is contained in game master’s memory and the values of grounded variables...
-
[33]
Like agents, the game master has an associative memory implemented using various components
The game master is implemented in a similar fashion to a generative agent. Like agents, the game master has an associative memory implemented using various components. However, instead of contextualising action selection, the components of the game master describe the state of...
-
[34]
| f e(zt, at)) (14)
The game master generates an event statement et in response to each agent action: et ∼ p(. | f e(zt, at)) (14)
-
[35]
After adding the event statement et to its memory the game master can update its components using the same Eq
The above equation highlights the fact that the game master generates an event statementet in response to every action of any agent, while the agent might take in several observations before it acts (or none at all). After adding the event statement et to its memory the game m...
-
[36]
In case the game master judges that a player did not observe the event, no observation is emitted. Notice that the components can have their internal logic written using any existing modelling tools (ODE, graphical models, finite state machines, etc.) and therefore can bring k...
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.