{"id":"76f73742-e872-48e1-b1cf-01c1133a7428","arxiv_id":"2411.17724","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"In simulated micro-economies, a semi-libertarian/utilitarian voting government and an inclusive institution yield higher ratios of house building to trading under both MARL and GABM, but the evidence base is a single run per setting.","lead":"This paper compares two AI simulation approaches, reinforcement learning and large-language-model agents, to see how different government styles change whether agents build houses, trade houses, or trade building skills in a small virtual economy. The findings suggest that a voting-based semi-libertarian government and an inclusive institution push agents to build rather than trade, but the simulations were run only once per setting, so the results are preliminary.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MARL ordering relies on pooling runs with different planner reward functions; with one run per configuration and no error bars, the Semi-Libertarian/Utilitarian and Inclusive ranking may be an artifact of pooling.","rationale":"Agree with the reader: the load-bearing assumption is statistical informativeness of the pooled runs. The paper is transparent and shares code, and the simulation design is plausible, but the central comparative claim lives or dies on Figures 4 and 5. Because each plotted value averages exactly two runs that differ in the planner's reward function, and the paper admits one simulation per parameter set, the design cannot separate governance-system effects from reward-function effects or seed noise. Since the AI-Economist planner's reward function directly shapes tax policy and hence labor/build/trade incentives, this is a substantive confound, not a cosmetic issue. A feasible reanalysis—disaggregating by reward function and adding seed replicates—would either confirm the ordering per reward function or show it is an artifact. The existing CONDITIONAL verdict already captures the uncertainty; this stress-test does not move it.","tokens_in":21790,"tokens_out":4531,"duration_ms":45225,"concrete_test":"Recompute Figs. 4 and 5 from the individual runs in Fig. 15 without pooling: for each governing system/institution, plot the build-to-trade-house and build-to-trade-skill ratios separately for each planner reward function, and repeat with at least 10 random seeds per condition. Check whether Semi-Libertarian/Utilitarian and Inclusive are strictly higher than the alternatives within each reward function and whether bootstrap or nonparametric intervals overlap. If the ordering appears under only one reward function or disappears under seed variation, the headline claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Figures 4 and 5, which support the abstract's headline claim for the extended AI-Economist, are computed by averaging over two simulations per condition that differ in the central planner's reward function (Fig. 15), and Section 4 states that the number of simulations per parameter set is one. The planner's reward function determines tax policy, and agents' build/trade/skill decisions respond to taxes; the two pooled runs are therefore not interchangeable replicates but different experimental conditions. Pooling them can create the appearance that Semi-Libertarian/Utilitarian and Inclusive produce the highest build/trade ratios even if neither reward function individually supports that ordering. There are no seed replicates, no error bars, and no statistical comparison, so the observed rank could flip under either reward function or under random seed variation. The Concordia results (Figs. 10-12) have the same single-run/qualitative limitation and are described with hedged language. The concern is not that the simulations are invalid, but that the current data do not establish the governance-ranking claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the AI-Economist MARL framework and the Concordia GABM framework so that agents can build houses, trade houses, and trade house building skill, and it introduces three governing systems (Full-Libertarian, Semi-Libertarian/Utilitarian, Full-Utilitarian) and, in the MARL variant, three governing institutions (Inclusive, Arbitrary, Extractive). The headline claims are that in the extended AI-Economist, the Semi-Libertarian/Utilitarian system and the Inclusive institution produce the highest ratios of building to trading houses and of building to trading skill (Figs. 4-5), and that in the extended Concordia, Full-Utilitarian produces more building and skill trading when the game master maximizes equality, while Full-Libertarian produces more of all three activities when it maximizes productivity (Figs. 11-12). The paper also discusses correlations between economic activities and welfare measures and offers a qualitative comparison of MARL and GABM capabilities, while explicitly acknowledging that only one simulation per parameter set was run.","tokens_in":21978,"tokens_out":3499,"duration_ms":35674,"significance":"If the headline results held, the paper would offer an interesting cross-framework comparison of how governance shapes economic incentives in simulated societies, and it would be a useful exploratory contribution with practical value for in-silico social science. The author provides open-source code for both frameworks, which is a genuine strength and enables reproducibility. However, the statistical foundation of the central comparisons is too weak: with a single simulation per configuration, no seed replicates, no error bars, and pooling of runs that differ in the planner's reward function, the ordering of governing systems is not established. The paper itself acknowledges this limitation, but the abstract and results sections still present the rankings as findings, which overstates the evidence. The GABM results are described in deliberately hedged language and are based on visual inspection, further reducing the strength of the second main claim.","major_comments":[{"comment":"The central MARL claim—that Semi-Libertarian/Utilitarian and Inclusive produce higher build/trade ratios—is computed by averaging over two runs per condition that differ in the central planner's reward function (Fig. 15 caption, Figs. 4-5). The author explicitly states in §4 that 'the number of simulations for each set of input parameters ... is one,' so these two runs are not replicates but distinct experimental conditions. Because the planner's reward function determines tax policy and agents' build/trade decisions respond to taxes, pooling such runs can create a spurious ordering even if neither reward function individually supports it. This is a load-bearing issue: without separating these conditions or adding proper replicates, the headline ranking is not supported by the presented data.","section":"§3, Figs. 4-5, Fig. 15, §4 Current Limitations"},{"comment":"No seed replicates, error bars, or significance tests are provided for any condition in either framework. With a single trajectory per parameter set in the MARL experiments, the observed ratios could flip under a different random seed; the paper notes that two-level RL training is 'particularly unstable' (Fig. 16 caption), which makes the absence of multiple runs especially concerning. For the GABM results, claims are made from visual inspection of ten episodes, with the author using hedged language ('it seems') and reporting one 'hallucination' (Fig. 10). The lack of any uncertainty quantification means that even the qualitative ordering of governance systems is unverified.","section":"§3, Figs. 4-5, 11-12"},{"comment":"The Concordia results constitute a second main claim of the abstract, yet they are reported only qualitatively. Figure 11 is interpreted as showing that Full-Utilitarian leads to 'slightly higher' building and skill trading under equality, and Figure 12 as showing that Full-Libertarian leads to higher activity under productivity, but no quantitative measure, statistical comparison, or sensitivity analysis is given. This is not a presentation issue but a limitation of the evidence for a central claim: the paper's own description ('it seems') acknowledges that the results are not robustly demonstrated.","section":"§2.2, §3, Figs. 10-12"}],"minor_comments":[{"comment":"The word 'fro' appears in 'fro both MARL and GABM'; a typo for 'for'. Several figure captions also contain grammatical errors, e.g., 'At it is clear form these plots' (Figs. 11-12) and 'refering' (Fig. 7).","section":"§4 Conclusions"},{"comment":"The text refers to 'the four resources' when describing the Semi-Libertarian/Utilitarian planner's investment of tax revenue, but the environment has only three material resources (wood, stone, and iron); this is inconsistent with the rest of the paper and should be corrected to 'three resources.'","section":"§2.1"},{"comment":"The phrase 'somewhat similar worlds' is appropriate given the differences between the MARL and GABM implementations, but the paper should explicitly state that the two frameworks are not directly comparable in a quantitative sense, since one uses learned policies and the other uses LLM prompting. The current language in the Introduction ('as similar as possible') is still vague; a short paragraph detailing the key differences (e.g., learning dynamics, action spaces, observation grounding) would strengthen the comparison.","section":"Abstract and §1"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its limitations and provides reusable code, which is commendable. However, the statistical issues are substantial and go beyond presentation: the main rankings are built on pooled runs with different planner rewards and no error bars. A major revision would need to either re-run the MARL experiments with multiple seeds per condition and separate analyses per reward function, or substantially reframe the claims as preliminary observations without ordering statements. Given that the Concordia results are also qualitative, the current evidence does not support the abstract's definite assertions. The heavy citation of the author's own prior work is not inherently problematic, but those works are not externally validated, which further limits the empirical grounding of the framework design choices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: the headline governance ranking in the MARL half is not statistically supported. The build-to-trade ratios that put Semi-Libertarian/Utilitarian and Inclusive on top come from pooling two runs per condition that differ in the planner's reward function, with one run per exact config and no seed replicates. That is a real confound, and the paper's own limitations section admits it.\n\nWhat's actually new: the direct comparison of extended AI-Economist and extended Concordia on the same governance question, with real extensions (three resources, three house types, house trading, skill trading, voting, tax component). That comparison is not in the prior literature, and the code is available. The author is also honest: the GABM results are explicitly hedged, and the limitations are stated plainly.\n\nSoft spots, in proportion: the pooling issue is load-bearing. The two pooled runs differ in the reward function of the central planner, which determines tax policy, which directly shapes agent build/trade decisions. They are not interchangeable replicates. The Concordia results have the same single-run limitation and are qualitative. So the abstract's specific ordering should not be trusted as a finding. The correlations in Figs. 8-9 also rest on the same pooled data. That said, this is an exploratory proof-of-concept, and the framing makes that clear.\n\nBottom line: worth sending to review because the comparison is novel, the code is shared, and the flaws are addressable, but it needs major revision before the claims can be taken seriously. I wouldn't cite the governance ranking in my own work yet; I might cite the methodological comparison if the reproducibility improves. The author is a serious thinker; the paper is coherent and honest on its own terms.","headline":"A genuinely novel MARL-vs-GABM comparison undermined by single-run pooling, but honest and worth a thorough revision.","tokens_in":22525,"tokens_out":2202,"would_cite":false,"duration_ms":21914,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A semi-libertarian, vote-counting government with an inclusive institution produces the highest build-to-trade ratios in a simulated economy; in an LLM-based world, Full-Utilitarian wins under equality goals and Full-Libertarian under…","keywords":["multi-agent reinforcement learning","generative agent-based modeling","governing systems","inclusive institutions","house building and trading","voting mechanisms","AI-Economist","Concordia"],"falsifier":"Run multiple independent seeds for each governing system and institution in both frameworks while holding the government's objective fixed, then recompute the build-to-trade and build-to-skill-trade ratios; if the Semi-Libertarian/Utilitarian and Inclusive ordering does not persist across seeds, the central claim is settled.","tokens_in":21539,"feed_emoji":"🏗️","tokens_out":6878,"duration_ms":65568,"temperature":0.7,"pith_summary":"This paper claims that in simulated economies where agents can build houses, trade houses, or trade house building skill, the governing system measurably changes which activity agents favor. In an extended multi-agent reinforcement learning (MARL) world, the Semi-Libertarian/Utilitarian system, a democratic-style arrangement in which agents vote and the planner implements their rankings, produces the highest ratios of building houses to trading houses and of building houses to trading house building skill; within that system, the Inclusive institution outperforms the Arbitrary and Extractive ones. In an extended generative agent-based model (GABM) using LLM-driven agents, the best system flips with the government's objective: an equality-focused game master makes Full-Utilitarian agents build more and trade skill more, while a productivity-focused game master makes Full-Libertarian agents do more building, house trading, and skill trading. The paper's purpose is to show that both MARL and GABM can simulate the same social phenomenon and that agents in both approaches partially infer the implicit rules of their world for planning.","feed_headline":"Voting-based governance boosts building over trading in simulated economies","feed_subtitle":"A democratic-style system and an inclusive institution raised build-to-trade ratios across two AI frameworks.","key_machinery":"The load-bearing machinery is the extended AI-Economist, a two-level deep MARL setup in which a central planner sets tax rates and invests tax revenue while mobile agents vote on resource priorities and choose among building houses, trading houses, and trading skill; and the extended Concordia, a generative agent-based model in which LLM-driven agents act through natural language and a game master translates actions, tracks grounded variables, and sets taxes to maximize equality or productivity.","core_discovery":"The paper extends the AI-Economist MARL framework so six agents gather wood, stone, and iron, build three house types, trade houses, and trade house building skill, with expert and novice agents distinguished by payment multipliers. It also extends Concordia so LLM-driven agents perform the same activities under a game master that tracks inventory, skill, build, vote, and tax. The central discovery is that governance structure changes which activity agents favor: in the MARL world, the Semi-Libertarian/Utilitarian system, where agents vote and the planner follows their Borda ranking, gives the highest ratios of building houses to trading houses and to trading skill, and among its institutions Inclusive gives higher ratios than Arbitrary and Extractive. In the Concordia world, an equality-optimizing game master makes Full-Utilitarian agents build and trade skill more, while a productivity-optimizing game master makes Full-Libertarian agents build, trade houses, and trade skill more. The paper also reports that the three economic activities correlate positively with productivity, equality, and maximin, with one exception: building houses correlates negatively with equality.","pith_inferences":["The pooled two-reward-function design leaves open that the governance ranking is an artifact of the government's objective; disentangling the two reward functions in the plots would test this directly.","If the ranking survives a random-seed sweep, governance-system simulations could become a low-cost screen for institutional designs before field experiments.","The negative building-versus-equality correlation found in both frameworks suggests a possible trade-off between construction activity and equality that the paper does not develop."],"forward_implications":["Under the extended AI-Economist, the Semi-Libertarian/Utilitarian system, the closest to current democratic systems, yields the highest build-to-trade and build-to-skill-trade ratios among the three governing systems.","Within that governing system, the Inclusive institution yields higher build-to-trade and build-to-skill-trade ratios than the Arbitrary and Extractive institutions.","In the extended Concordia, an equality-optimizing game master makes Full-Utilitarian agents build more houses and trade more house building skill, while a productivity-optimizing game master makes Full-Libertarian agents build more houses, trade more houses, and trade more skill.","In both frameworks, building, trading houses, and trading skill mostly correlate positively with productivity, equality, and maximin, except that building houses correlates negatively with equality.","Both MARL and GABM agents partially infer the implicit rules of the environment, suggesting both approaches can model similar social phenomena despite their different architectures."],"supporting_citations":[{"why":"Supplies the original AI-Economist two-level MARL framework and formalisms for agents, planner, taxation, and social welfare that the paper extends with trading, voting, and skill.","marker":"Zheng et al. (2022)"},{"why":"Supplies the original Concordia GABM library and the game-master and generative-agent architecture that the paper extends with inventory, skill, build, vote, and tax tracking.","marker":"Vezhnevets et al. (2023)"},{"why":"Established the voting-based governing systems along individualistic-collectivistic and discriminative axes that this paper reuses for its governance comparisons.","marker":"Dizaji (2023b)"},{"why":"Provides the MARL convergence and equilibrium-selection caveats that justify the 5000-step training limit and the stated limitation on interpreting the MARL results.","marker":"Albrecht et al. (2023)"}],"fun_headline_variants":["Governance shifts AI agents from trading to building houses","Voting-based rule boosts house building over trading in AI worlds","Inclusive institutions raise build-to-trade in simulated economies","AI simulations show governance changes economic focus","Equality vs productivity alters AI house-building behavior"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is statistical: because each configuration was run once and the displayed ratios pool two runs that differ in the government's objective, the ranking of governing systems could flip if the reward difference or the single draw of randomness is responsible for the pattern.","fun_headline_variants_meta":{"raw":{"variants":["Governance shifts AI agents from trading to building houses","Voting-based rule boosts house building over trading in AI worlds","Inclusive institutions raise build-to-trade in simulated economies","AI simulations show governance changes economic focus","Equality vs productivity alters AI house-building behavior"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1544,"prompt_tokens":1145,"completion_tokens":399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":761,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":761,"tokens_out":399,"duration_ms":4128,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:42:26.884199+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run multiple independent seeds for each governing system and institution in both frameworks while holding the government's objective fixed, then recompute the build-to-trade and build-to-skill-trade ratios; if the Semi-Libertarian/Utilitarian and Inclusive ordering does not persist across seeds, the central claim is settled.","supporting_citations":[],"review_version":1}