{"id":"8d260d38-84ea-4193-95fb-e37657fa2dc9","arxiv_id":"2602.10429","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":10,"one_line_summary":"A deployed LLM-agent society with branch-thinking planning and dual-process memory claims to reproduce heavy-tailed returns, volatility clustering, and education-driven wealth stratification, but key evidence reduces to hand-configured rules without released code or error bars.","lead":"AIvilization v0 is a public, large-scale artificial society where LLM-driven agents work, trade, study, and sleep under hard resource constraints, with humans steering agents by setting goals or issuing commands. The paper reports real-economy-like market patterns and wealth gaps, but the evidence is thin and some headline results are partly hardwired into the rules.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Market stylized facts may be a zero-inflation artifact: fish price span is 0.001 in log terms, yet Table 1 reports kurtosis ~9.5 and |r| ACF ~0.19; no null model rules out discretization.","rationale":"I focused on the market stylized facts because they are the strongest external-validation claim and the least protected by construction. The stratification result is less risky as evidence: Section 4.4 explicitly says the link between occupations and wealth is written into the rules, so even if the sorting dynamics are interesting, the 'driven by education and access constraints' part is partly by design. The ablation, by contrast, supports an architecture claim; it lacks variance estimates and in Task 2 the Without-OD variant slightly beats Default, which the paper explains away. But the abstract's 'research-grade artificial society' claim leans on the market regularities, and that is where a single artifact can overturn the conclusion.\n\nThe reader's weakest assumption identifies the same spot: near-constant fish prices with high kurtosis. My concern sharpens it: the test should be a null model of price-grid quantization and time-varying trade intensity. The paper's own Section 7 admits the stratification analysis is observational and that LLM hallucinations/constraints remain, but it does not flag the discretization confound in Section 4.3. The lack of code/data makes it impossible to audit the OHLC construction. Therefore rejection remains appropriate; if the null model fails to reproduce the statistics, the paper could become a CONDITIONAL accept for the platform architecture with weaker claims.","tokens_in":23944,"tokens_out":6806,"duration_ms":81842,"concrete_test":"Recompute the fish-market statistics from raw tick data. First, report the minimum nonzero log-return (the AMM tick size) and the fraction of 5-minute bins with zero return. Then construct a placebo: keep the empirical trade arrival times and pool depths, but assign each trade's direction (buy/sell) and size by independent draws from the empirical joint distribution (or shuffle signs across trades); generate the resulting OHLC series and recompute excess kurtosis, skewness, ACF(|r|), and Ljung-Box p for fish. If the placebo reproduces Table 1's fish row within sampling error, the stylized-fact claim is an artifact. Cross-check by computing the same statistics on event-time returns restricted to nonzero returns; if kurtosis and ACF collapse, zero-inflation is the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing empirical claim is that the deployed market 'reproduces key stylized facts (heavy-tailed returns and volatility clustering)' (abstract; Section 4.3). This claim rests entirely on 5-minute OHLC series constructed from AMM transactions. The fish series used as the lead case has log-price range 0.001 and max drawdown 0.0715% (Section 4.2, Figure 4); the price is effectively pinned to a tiny discrete grid. Under such quantization, most intervals contain a zero return and the remaining returns are one-tick jumps. The standardized distribution then has a large spike at zero and thin-looking tails beyond the tick, mechanically generating the reported excess kurtosis (9.489 for fish) without any economic mechanism. The ACF of |r| (0.189 for fish) and the Ljung-Box rejection can likewise be produced by any time-varying trade intensity (volume clusters, day/night cycles, etc.), which is not volatility clustering in the finance sense and does not require agent-level feedback. The paper provides no null model, no minimum-price-increment analysis, no event-time returns, and no raw trade-size/pool-depth context. So Table 1 does not discriminate between endogenous market dynamics and a discretization/activity artifact. Because this is the central external-validity evidence, the claim that the platform is a research-grade artificial society loses its empirical foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AIvilization v0, a publicly deployed large-scale artificial society coupling an LLM-agent architecture (branch-thinking planner, dual-process memory, human-in-the-loop steering) with a resource-constrained economic environment (physiological costs, multi-tier production, AMM pricing, gated education-occupation system). Using 400,000 transactions from a mature phase of the platform, the authors construct 5-minute OHLC series and report that simulated markets reproduce heavy-tailed returns and volatility clustering, and that wealth stratification is driven by education and access constraints. A controlled ablation study compares the full planner with two simplified variants on multi-objective and simple tasks. The central claim is that the platform is a research-grade artificial society for studying emergent macro-social phenomena.","tokens_in":24325,"tokens_out":2718,"duration_ms":34012,"significance":"If the empirical claims held, this would be a notable contribution: a large-scale, public LLM-agent society with tens of thousands of agents and 600k+ transactions, showing real-economy-like statistical regularities and structured inequality. The architecture itself—hierarchical branch planning, adaptive profiles, memory-mediated steering—contains useful ideas and the ablation study is a reasonable start. The paper ships no code or data release, and the central validation rests on one week of fish-market OHLC data. The claimed stylized facts are, however, plausibly artifacts of price discretization, and the stratification result is largely forced by the model's own eligibility equations. These issues are load-bearing for the paper's headline contribution, so the significance of the paper as it stands is limited.","major_comments":[{"comment":"The central claim that the market 'reproduces key stylized facts (heavy-tailed returns and volatility clustering)' is not supported because the reported statistics are consistent with a discretization artifact. The fish price is confined to [304.398, 304.808] (log range 0.001) with a maximum drawdown of 0.0715% (§4.2). Under such extreme quantization, returns are zero in most 5-minute bins and one-tick jumps otherwise; the standardized distribution automatically has a large spike at zero and excess kurtosis of 9.489, and the ACF of |r| can be produced by any time-varying trade intensity. The paper provides no null model, no minimum-tick analysis, no event-time returns, and no pool-depth/trade-size context. Table 1 therefore does not discriminate between endogenous economic dynamics and a zero-inflation artifact.","section":"§4.2, §4.3, Table 1, Eqs. (18)–(19)"},{"comment":"The claimed 'emergent' wealth stratification is to a large degree hard-wired by the model's own rules. Equation (10) gates occupation eligibility on dynamic knowledge thresholds; Eq. (12) defines these thresholds as the (1−πj) quantile of the education distribution, so eligibility shares are fixed by hand-chosen πj parameters (Table 11); Eq. (15) makes high-tier wages increase with the very same threshold. Since wealth is tied to occupation wages, the positive education-wealth gradient and occupation-tier wealth ordering shown in Figures 9–10 follow almost algebraically. The authors acknowledge in §4.4 that stratification is 'an outcome of the simulation’s core rules,' but they nevertheless claim a dynamic sorting process. No counterfactual (e.g., fixed thresholds, no residential gates, or random πj) is provided to separate the rule-forced component from truly emergent sorting. Thus the","section":"§4.4, Eqs. (10)–(12), (15), Table 11"},{"comment":"The ablation conclusions are overstated relative to the evidence. In Task 2 (Table 3), the Default planner ranks second on both net worth and education score, behind Without-OD; the paper acknowledges this but the overall discussion (§5.4) claims the full architecture 'consistently' and 'materially' outperforms on complex multi-objective tasks. No standard errors, confidence intervals, or multiple seeds are reported—each condition uses 80 agents with no indication of replication. The differences (e.g., Task 1 net worth 110,098 vs. 95,279) may be real, but with a single run the claim of robustness is not statistically grounded. At minimum, the paper should present per-agent distributions and effect sizes, not only means.","section":"§5, Tables 2–5"}],"minor_comments":[{"comment":"Typo: 'Their Their long-term goal' should read 'Their long-term goal.'","section":"§5.1"},{"comment":"Typo: 'econciling' should be 'reconciling.'","section":"§2.1"},{"comment":"The table reports no sample size, number of 5-minute intervals, or time span per commodity; with 400,000 transactions split across ten assets, intervals may vary widely. Also, p-values are shown as '<10⁻⁶' without reporting the Ljung-Box statistic or the number of lags used.","section":"Table 1"},{"comment":"The deployment data uses a 7× time compression, while the ablation uses 35×. The paper should clarify whether the 35× scaling is used in the mature-phase transaction data or only in the controlled experiments, since this affects the interpretation of 5-minute returns.","section":"§4.1 and §5.1"},{"comment":"The price series is visually displayed on an extremely narrow y-axis (304.398–304.808). The authors should also show the corresponding number of trades per bin and the pool depth to allow readers to judge whether the 'micro-fluctuations' are genuine price discovery or just quote noise.","section":"§4.2, Figure 4"}],"recommendation":"reject","confidential_remarks":"The paper is a system/architecture paper with a strong empirical claim. The architecture and deployment are interesting, but the market-validation section does not survive basic scrutiny: the fish price is pinned to a 0.001 log range, so the reported kurtosis and volatility-clustering statistics are almost certainly discretization artifacts. The stratification analysis is also largely circular because eligibility shares and wage equations are chosen to produce the observed hierarchy. These are not presentation issues—they undermine the central claims. The ablation study is also under-powered. I would encourage the authors to resubmit with a proper null model for the market statistics, coarser or event-time returns, counterfactuals for the gating rules, and replicated ablations with distributions. In its current form, I cannot recommend acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read this with some anticipation because the platform is genuinely ambitious. The engineering is real: a public deployment with tens of thousands of LLM agents, a unified planner-memory-steering loop, a coupled AMM economy with tiered production and gated education-occupation, and a high-frequency transaction log of over 600k records. That is not a toy. The ablation study is small but goes after the right question, and the authors are honest that simpler planners match the full one on narrow tasks. Section 7's limitations are candid—observational stratification, LLM hallucination, compute bottlenecks. That honesty counts.\n\nThe market validation is the weak link. The fish market, their lead case, has a log-price range of 0.001 and a max drawdown of 0.07%. That price is essentially pinned to a discrete tick. On such a series, most intervals have zero return and the occasional trade moves the price one tick, which mechanically produces the reported excess kurtosis around 9.5 and a significant ACF of |r|. Without a null model, tick-size analysis, or event-time returns, Table 1 does not discriminate between endogenous dynamics and a discretization artifact. The Ljung-Box test on absolute returns mostly detects time-varying trade intensity, which is not volatility clustering in the usual finance sense.\n\nThe stratification result is also not emergent. Section 4.4 admits it is an outcome of the core rules, and the hand-set eligibility shares and wage equation (Eqs. 10-12, 15) directly determine the occupational wealth gradient. The conclusion calls it \"emergent formation,\" which is stronger than the body supports. The ablation is single-run, no variance, and Task 2 shows the full planner underperforming Without-OD; the explanation is plausible but the evidence is thin. No code or data is released, which is a problem when the main claims are empirical.\n\nWho gets value from this? People working on LLM-agent social simulation and game-like testbeds will get a useful systems overview and a cautionary example of validation pitfalls. It deserves a serious referee, because the platform is real and the questions matter. But the referee should demand a null model for the market data, a tick-size robustness check, multi-seed ablations, and code/data release. As submitted, the strong claims in the abstract and conclusion are not supported by the evidence.","headline":"Real platform, real engineering, but the numbers don't back the headline claims: the market 'stylized facts' are likely discretization artifacts and the stratification is mostly rule-driven.","tokens_in":24843,"tokens_out":2603,"would_cite":false,"duration_ms":29494,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AIvilization v0 claims that a publicly deployed society of tens of thousands of LLM-driven agents can sustain long-horizon autonomy and, in its mature phase, generate markets whose returns are heavy-tailed and volatility-clustered plus educ","keywords":["artificial society","LLM agents","multi-agent simulation","stylized facts","volatility clustering","heavy-tailed returns","wealth stratification","automated market maker"],"falsifier":"Take the same 400,000-transaction block, rebuild prices at coarser granularity (e.g., 30-minute or 1-hour bins), and exclude intervals with zero trades or identical quotes; if excess kurtosis collapses toward Gaussian levels and the Ljung-Box test no longer rejects independence of |r|, then the reported stylized facts are artifacts of binning and price discretization.","tokens_in":23791,"feed_emoji":"📈","tokens_out":4222,"duration_ms":45761,"temperature":0.7,"pith_summary":"This paper is trying to establish that a large, publicly deployed population of LLM agents—steered by humans and operating under hard resource constraints—can form a coherent artificial society rather than collapsing into chaos or trivial loops. Its central claim is that the market data from the platform's mature phase reproduces two canonical financial stylized facts (heavy-tailed returns and volatility clustering) and that wealth stratification emerges from the interaction of education investment and residential-tier barriers. If true, this would make the platform a research-grade testbed for studying emergent macro-social phenomena from micro-level agent decisions. The paper also argues that its hierarchical Branch-Thinking Planner, paired with dual-process memory, is necessary for reliable long-horizon multi-objective behavior, while simpler planners suffice for narrow tasks.","feed_headline":"AI agent society reproduces heavy-tailed markets and wealth gaps","feed_subtitle":"Tens of thousands of LLM agents in a live sandbox produced volatility clustering and education-driven stratification, the authors report.","key_machinery":"The load-bearing machinery is a three-part loop: (1) a Branch-Thinking Planner that decomposes a life goal into parallel objective branches and uses context-based prioritization plus pre-execution Action Simulator rollouts to keep actions feasible; (2) a dual-process memory that separates short-term execution traces from long-term semantic consolidation, letting identity persist yet evolve; and (3) a constant-product Automated Market Maker (IS_i·CR_i=k) that sets prices through liquidity-pool ratios and couples money supply to real output. The AMM is what converts agent actions into a price series, and the education-occupation gate is what converts human-capital investment into wage and weal","core_discovery":"On the paper's own terms, the discovery is that a persistent simulated economy built from LLM agents—who buy, sell, produce, study, and sleep under physiological and eligibility constraints—spontaneously displays real-world market regularities. In a block of 400,000 high-frequency transactions from the mature public deployment, the fish market's price stayed between 304.398 and 304.808 (log-price range 0.001), yet 5-minute log returns across ten commodities show excess kurtosis above 6 (up to 9.873) and significant lag-1 autocorrelation of absolute returns, which the authors interpret as heavy tails and volatility clustering. The same logs show a monotonically increasing, nonlinear relation","pith_inferences":["The stylized-fact evidence rests on a single week of one liquid market; a natural extension would check whether the same excess kurtosis appears in the silicon and wood supply chains' price series while controlling for tick size.","If the fish price barely moved, the heavy tails may reflect the discreteness of 5-minute bins or AMM pool granularity; a cleaner test would use trade-level returns or compare against a null model of random trades through the same AMM.","The paper leaves implicit that the same architecture could be used to study institutional changes, such as removing residential barriers, and measure their effect on inequality; an A/B experiment varying the education threshold quantile is a concrete next step.","Because the deployed population includes human-steered agents, the data confounds autonomous emergence with human guidance; the causal claim about steering would need a fully autonomous control cohort."],"forward_implications":["If the market regularities are genuine, then agent-based social simulators with hard constraints and LLM decision-making can produce emergent financial statistics without being explicitly calibrated to do so.","The platform's wage regime, tying dynamic wages to a knowledge-threshold quantile, implies that rising average education will automatically raise entry requirements for top jobs, preserving positional competition as the population upskills.","The ablation results imply that a lightweight planning route is enough for simple tasks, so future systems can save compute by activating the full branch-thinking stack only for multi-objective, long-horizon goals.","The correlation between early educational steering and upward mobility, if causal, suggests targeted long-horizon prompts could be a policy lever inside such simulations.","The architecture's memory-mediated steering suggests a path to hybrid-autonomy platforms where human influence is absorbed into agent identity rather than overwriting prompts."],"fun_headline_variants":["AI agent economy spontaneously produces fat-tailed returns","Tens of thousands of LLM agents yield realistic market stats","Simulated society of AI agents reproduces wealth stratification","LLM-powered virtual economy exhibits volatility clustering","AI agent city: stable prices, fat tails, and class divides"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The empirical validation treats one week of 5-minute OHLC data from the fish market—where prices move over a range of only 0.001 in log space—as a meaningful price-discovery series; if those tiny movements are rounding artifacts or pool-granularity noise rather than endogenous economic dynamics, the headline stylized-facts result loses its evidentiary base.","fun_headline_variants_meta":{"raw":{"variants":["AI agent economy spontaneously produces fat-tailed returns","Tens of thousands of LLM agents yield realistic market stats","Simulated society of AI agents reproduces wealth stratification","LLM-powered virtual economy exhibits volatility clustering","AI agent city: stable prices, fat tails, and class divides"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1811,"prompt_tokens":829,"completion_tokens":982,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":905}},"tokens_in":573,"tokens_out":982,"duration_ms":10777,"temperature":1.0,"reasoning_tokens":905,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T01:16:02.751863+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 400,000-transaction block, rebuild prices at coarser granularity (e.g., 30-minute or 1-hour bins), and exclude intervals with zero trades or identical quotes; if excess kurtosis collapses toward Gaussian levels and the Ljung-Box test no longer rejects independence of |r|, then the reported stylized facts are artifacts of binning and price discretization.","supporting_citations":[],"review_version":1}