{"id":"98a13051-e581-40f0-a848-c109700073ae","arxiv_id":"2608.10262","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Four Chinese frontier LLMs show lab-level divergence in aggressive-equilibrium rates (P_A 1% to 9%), with within-ecosystem spread exceeding the East-West mean difference.","lead":"This paper tests four Chinese AI models in a repeated cooperation game, using the same translator for all of them so the comparison is fair. It finds the four labs differ more from each other than the Chinese group differs from Western models, so 'Chinese models' should not be treated as one bloc.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"H6's lab-level claim rests on one model per lab and a P_A contrast that is partly driven by the near-zero lower bound; the internal statistics support divergence, but the 'lab not model' inference is underdetermined.","rationale":"The reader's verdict is CONDITIONAL, and their weakest assumption was that the lab-level inference assumes a single flagship model represents each lab's alignment lineage. I agree with that identification; it is the most load-bearing concern. The paper itself flags it in Section 5.4 ('the panel is four labs at one snapshot in time; model versions drift') and also flags the converter difference in the Western baseline. Under the reviewing rules, I must flag those self-identified limitations explicitly, and they do bound the central claim. The paper's strongest internally supported result is the H6 pairwise divergence among the four Chinese models under a fixed converter. The statistical analysis is appropriate: Holm-Bonferroni correction, n=500, robustness check with a second converter that preserves the P_A ordering. However, the central claim as framed—'the lab, not the ecosystem, is the unit at which cooperative disposition is set'—goes beyond what the design can support. With one model per lab and no within-lab replication, the observed P_A differences could be model-version effects, serving-backend effects, or even random variation in what is effectively a sample of four model instances. The paper does not inflate this; it is explicit about the limitation. But the headline conclusion is, as stated, underdetermined by the data. The question is whether this warrants changing the verdict. The reader already made the verdict CONDITIONAL. My stress-test does not identify an additional flaw that would make the paper incorrect in its internal logic; the internal argument is honest and the limitation is acknowledged. It is a paper that would require a Western fixed-converter re-run and ideally a second model per lab to become ACCEPT, but the current verdict CONDITIONAL already captures exactly that. Therefore I maintain UNCHANGED relative to the reader's verdict. The alternative would be to move to CONDITIONAL if I were starting from ACCEPT, but the verdict already reflects the caveat. I see no reason to REJECT: the H6 result is internally consistent, the statistics are proper, and the robustness check is meaningful. My concern is about the strength of the inference from single models to labs, which is a known and self-acknowledged limitation, not an internal inconsistency. The concrete test I propose—re-running with a second model per lab—is the minimal experiment that would settle it. If the P_A ordering and pairwise significances survive, the lab-level claim would be substantially strengthened. If not, the paper's headline overreaches. This is the single most load-bearing check because the entire 'not a monolith' claim and the design implication in Section 5.3 (provenance at the lab level matters) depend on the lab being the unit, not the model or serving configuration. The paper's internal evidence for the H6 pairwise differences among those four specific model instances is solid; the jump to labs is the soft spot.","tokens_in":13726,"tokens_out":2513,"duration_ms":22235,"concrete_test":"Re-run H6 with a second model from each of the four labs (e.g., the previous flagship or the current smaller model, at n=500 per condition, same fixed converter). If the P_A ordering and the four significant pairwise differences do not survive the within-lab model swap, the inference that the lab, not the model, sets cooperative disposition would not be supported. This is exactly the experiment the paper's Section 5.4 admits is missing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of this paper is H6: the four Chinese labs differ significantly in aggressive-equilibrium proportion P_A, and this is interpreted as lab-level divergence in cooperative disposition. The statistics are internally sound; the concern is whether the design supports the causal and interpretive load placed on them. The design samples exactly one served flagship model per lab, at one point in time, via one gateway (with serving backend and quantization uncontrolled). The 8pp P_A spread (1-9%) is real against n=500 sampling error, but the four observations are also just four models; the inference that the lab is the unit of behavior is an extrapolation from a single model per lab. The paper explicitly acknowledges this in Section 5.4 and the reader flagged it. This is a genuine limitation for the central claim. In addition, the 'within-ecosystem variation exceeds the East-West gap' headline compares the P_A spread of four Chinese labs (SD 4.1pp) with the difference of the two ecosystem means (5.0% vs 5.0%). That comparison is a literature contrast where the Western P_A values were measured under a different, per-provider converter. The paper carefully labels this as a literature contrast (Section 5.2, Table 5 footnote), so it is not an internal inconsistency, but the headline conclusion depends on a measurement whose comparability is not established. The stronger internal result is the four significant pairwise comparisons among the Chinese labs; that is well supported. The load-bearing but under-supported step is the final interpretive move from 'four single models differ' to 'the lab is the unit of behavioral variation and Chinese models are not a monolith.' A single model per lab cannot rule out model-version or serving-configuration effects.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the cooperative versus aggressive equilibria of four Chinese frontier LLM agents (DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5, GLM-5.1) in an evolutionary Iterated Prisoner's Dilemma. To remove a confound present in prior work, the author holds the natural-language-to-code converter fixed (GPT-5.4 Mini) across all labs, so that cross-lab differences in equilibrium outcomes cannot be attributed to coding ability. The study evaluates two pre-registered hypotheses: H5, that the cooperative-plurality bias generalizes to Chinese models, and H6, that Chinese-model behavior is not monolithic. The paper reports qualified support for H5 (6 of 12 lab-prompt combinations show a cooperative plurality, versus 9 of 12 in the Western baseline, with the difference not statistically significant and sensitive to the converter) and support for H6: pairwise z-tests on the aggressive-equilibrium proportion P_A in the balanced noiseless Default condition yield four of six significant comparisons after Holm-Bonferroni correction, with P_A ranging from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro. The paper concludes that the lab, not the ecosystem, is the unit at which cooperative disposition is set. A replication package containing the strategy libraries, equilibria, and code is released.","tokens_in":14027,"tokens_out":7744,"duration_ms":79851,"significance":"If the main result holds, this is a worthwhile contribution to the empirical study of LLM cooperation. The fixed-converter design is a genuine methodological improvement over provider-aligned conversion, and the H6 statistical analysis is appropriately conservative: six pairwise comparisons with Holm-Bonferroni correction on n=500 runs per condition. The paper also ships a full replication package, which supports reproducibility. I find no circularity problem in the central measurement: P_A is read off Moran-process simulations, not fitted to produce the result. The strength of the evidence for divergence among the four served models is high, and the finding that within-ecosystem behavioral variation is structured and large is practically relevant for multi-agent system deployment. The main weakness is interpretive: the paper moves from significant differences among four model instances to a claim about labs as the causal unit, and its headline East-West comparison rests on a literature contrast across different converters. These issues affect the framing and scope of the central claim but not the internal validity of the pairwise statistical test.","major_comments":[{"comment":"The central claim that 'the lab, not the ecosystem, is the unit at which cooperative disposition is set' is not fully supported by the design: each lab contributes exactly one served flagship model at a single point in time, accessed through one gateway with uncontrolled serving backend and quantization (Table 2; acknowledged in Section 5.4). The statistically significant pairwise z-tests therefore establish divergence among four model instances, not among four labs as a class. To carry the lab-level attribution, the paper would need multiple checkpoints or model versions per lab, or the conclusion should be explicitly restricted to the evaluated served flagship models. This is a load-bearing interpretive step for H6 and for the title.","section":"Section 5.2, Table 2, Section 5.4"},{"comment":"The headline that within-ecosystem variation exceeds the East-West gap compares the spread of the four Chinese labs' P_A (SD 4.1pp) with the difference between the Chinese and Western mean P_A (5.0% vs 5.0%), but the Western values come from a different, per-provider converter (Table 5 footnote; Section 5.2 footnote 1). The paper carefully labels this as a literature contrast, yet the abstract and conclusion present the comparison as a substantive finding. Because the two sides were measured under different conversion regimes, the apparent absence of an East-West gap could be an artifact of the converter difference. The claim should be downgraded to a tentative literature contrast, or the Western models should be re-run under the fixed converter before it is used as a headline result.","section":"Section 5.2, footnote 1, Table 5"},{"comment":"The converter-robustness check re-converts only 10% of each library (8 of 75 strategies) and reports that P_A moves by at most 4pp and that the H6 structure survives. This is a weak perturbation: a 10% re-conversion cannot rule out that a full re-conversion with a different converter would shift P_A values or even reorder the labs, and the check mainly acts on near-tie Cooperative/Neutral cells rather than on the aggressive-equilibrium axis. Since Sections 5.1 and 6 use this check to argue that the H6 clusters survive converter choice and that the neutral lean is within converter noise, the limited power of the 10% re-conversion should be stated explicitly and the claims scaled back accordingly.","section":"Section 4.7, Table 9"}],"minor_comments":[{"comment":"The manuscript repeatedly states that hypotheses were 'pre-registered' but gives no link, timestamp, or registration document; please provide the preregistration in the replication package or as supplementary material.","section":"Sections 3.2 and 3.7"},{"comment":"Table 6 reports Holm-Bonferroni-corrected significance symbols but not the adjusted p-values or the correction threshold; please report them so that readers can verify the correction.","section":"Table 6"},{"comment":"The two-proportion z-test for 6/12 versus 9/12 treats the twelve lab-prompt combinations as exchangeable observations even though they are nested within four labs and three prompt styles; since the paper already calls this a literature contrast, it would be clearer to omit the p-value or to state explicitly that the effective sample size is four labs.","section":"Section 4.3"},{"comment":"The term 'Phase 1' is used without a citation or definition; if it refers to a separate report, please cite it, and otherwise define it in Section 2.","section":"Sections 2 and 3.2"},{"comment":"The caption of Table 9 uses notation such as 'C→N' for plurality flips; please define this notation in the caption.","section":"Table 9"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe one thing to know: this is a genuinely useful confound-control paper, not a headline result. Its fixed-converter design (GPT-5.4 Mini translates all strategies) cleanly removes the coding-ability confound that has plagued cross-provider IPD comparisons. The H6 divergence among four Chinese labs — P_A from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm-Bonferroni — is statistically well-supported. The converter-robustness check is the most careful part of the paper: it shows the aggressive-equilibrium structure survives a 10% re-conversion, while the cooperative/neutral plurality count is honestly flagged as converter-sensitive. Credit also for pre-registration, n=500, explicit limitations, and a released replication package.\n\nThe soft spots are where the reader put them. One model per lab is the biggest. The Section 5.2 claim that the lab is the unit of behavior is an extrapolation from four served flagship models at one point in time, accessed through one gateway with backend and quantization uncontrolled. Four observations cannot rule out model-version or serving effects. The paper acknowledges this in Section 5.4, but the abstract and conclusion still carry the unit-level inference. The East-West comparison is also a literature contrast: Western P_A values came from a per-provider converter, so the \"within-ecosystem variation exceeds the East-West gap\" line is not a measured result. The authors label it as a literature contrast, so it is not misleading in the text, but it should not be the headline.\n\nH5's 6/12 vs 9/12 is fairly presented as consistent but qualified, and the z = -1.26, p = 0.21 supports that. The near-ties are real, and the robustness check flipping plurality in five cells is good reason not to treat the neutral lean as firm.\n\nBottom line: this deserves peer review. The internal statistics and design support divergence among the four sampled models; the unit-level and ecosystem-level framing needs either more models per lab, a Western fixed-converter re-run, or deliberately narrower claims. A serious referee should engage rather than desk-reject.","headline":"A genuinely useful fixed-converter confound-control paper whose internal H6 statistics are sound, but whose headline 'lab is the unit' claim is underdetermined by one model per lab and whose East-West comparison is only a literature contrast.","tokens_in":14628,"tokens_out":1722,"would_cite":true,"duration_ms":18558,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chinese frontier AI models are not a monolith: in an evolutionary prisoner's dilemma, four Chinese labs diverge significantly in aggressive equilibria, and within-ecosystem spread exceeds the East-West gap.","keywords":["iterated prisoner's dilemma","large language model agents","evolutionary game theory","Moran process","cooperative AI","Chinese frontier LLMs","lab-level divergence","fixed-converter protocol"],"falsifier":"Re-run the full fixed-converter protocol with two or more served models per Chinese lab under the same converter. If within-lab $P_A$ differences are as large as the between-lab spread, or if the lab ordering flips, the lab-level attribution fails.","tokens_in":13507,"feed_emoji":"🤖","tokens_out":6906,"duration_ms":60829,"temperature":0.7,"pith_summary":"The paper asks whether the cooperative bias observed in Western LLM agents extends to Chinese frontier models and whether those models should be treated as a single bloc. To answer, it runs an evolutionary Iterated Prisoner's Dilemma with four Chinese flagship models using a fixed converter, so coding ability cannot contaminate the comparison. The central finding is that the four labs differ significantly in aggressive-equilibrium proportion $P_A$, from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, and the spread across the four labs is larger than the difference between the Chinese and Western ecosystems' means. A cooperative plurality appears in 6 of 12 lab-prompt combinations, which the paper reports as consistent with the Western cooperative bias but qualified by a lean toward neutral equilibria that sits within converter noise. If the finding holds, agent deployment should be chosen at the level of the laboratory, not the region.","feed_headline":"Chinese AI labs are not one bloc, prisoner's dilemma test finds","feed_subtitle":"In evolved prisoner's dilemma games, aggressive equilibria range from 1% to 9% across four Chinese labs — wider than the East-West gap.","key_machinery":"The load-bearing mechanism is the fixed-converter protocol: each lab's natural-language strategies are turned into executable Python by the same converter, so every cross-lab comparison isolates strategy generation. This is coupled to the standard evolutionary machinery of an all-play-all IPD tournament followed by a Moran process, a finite-population selection model, run at $n=500$ per condition across three prompt styles and four population regimes. The derived quantity $P_A$, the proportion of Moran runs ending in an all-aggressive monoculture in the balanced noiseless Default condition, carries the H6 test, and its pairwise comparisons are Holm-Bonferroni corrected.","core_discovery":"The paper's discovery is that cooperative disposition in Chinese frontier LLM agents is set at the lab level, not the ecosystem level. In the balanced noiseless condition, $P_A$ runs from 1% (Qwen3-Max) to 9% (DeepSeek V4 Pro), and four of six pairwise comparisons survive Holm-Bonferroni correction, splitting the four labs into a takeover-resistant pair (Qwen, Kimi) and a takeover-prone pair (DeepSeek, GLM). Because all strategies were converted by one fixed converter, these differences cannot be attributed to coding ability. The between-ecosystem comparison is a literature contrast rather than a controlled experiment, but on this measure the Chinese and Western mean $P_A$ are both about 5.0%, while the Chinese labs span 8 percentage points, so within-ecosystem variation exceeds the East-West gap. The paper also reports that the cooperative-plurality bias generalizes in attenuated form, with 6 of 12 lab-prompt combinations favoring cooperation, but treats the cooperative-neutral balance as converter-sensitive rather than a firm regime difference.","pith_inferences":["If lab-level divergence is real, then a single flagship sample per lab is too thin: a fair test of a lab's alignment lineage would need multiple models or versions per lab under the same fixed converter, and within-lab variance would need to be smaller than between-lab variance.","The converter-sensitivity of the cooperative-neutral boundary suggests that part of the measured 'cooperative bias' may live in the translation step rather than only in the model; deliberately varying converter families could locate where cooperative disposition enters.","The East-West mean comparison is a literature contrast across different converters; rerunning Western models under the same fixed converter would turn the tie into a controlled test, and mixed-provider populations would show whether lab dispositions compose or collide."],"forward_implications":["Treating 'Chinese models' as a monolith is not supported: selecting an agent by region could unknowingly pick between a population that resists aggressive takeover (Qwen, Kimi) and one that yields to it (DeepSeek, GLM).","Cooperative bias does appear in a non-Western alignment lineage, so it is not unique to Western models; in the balanced noiseless condition no lab's clean cell shows aggressive dominance.","Because the converter was fixed, the observed lab-level divergence cannot be explained away as a coding-ability artifact; it is a property of what the models generate.","The exact cooperative-plurality count (6/12 versus the Western 9/12) is converter-sensitive, so claims about weaker Chinese cooperation should not be drawn from this design."],"supporting_citations":[{"why":"Supplies the generate-strategies-then-evolve IPD benchmark and the Western cooperative-bias baseline this study inherits and tests against.","marker":"[20]"},{"why":"Provides the open all-play-all IPD tournament implementation used to run matches.","marker":"[9]"},{"why":"Gives the finite-population selection model that generates the equilibrium proportions.","marker":"[12]"},{"why":"Connects finite-population selection to cooperation in the IPD, motivating the equilibrium analysis.","marker":"[13]"},{"why":"Defines the Self-Refine prompt style used as one of the three prompting conditions.","marker":"[11]"},{"why":"Supplies the action-noise mechanism used in the noisy conditions.","marker":"[21]"},{"why":"Founds the IPD evolution-of-cooperation framework and its payoff structure.","marker":"[3]"}],"fun_headline_variants":["Chinese AI labs split on cooperation, new test shows","Lab, not ecosystem, sets AI cooperation: Chinese models diverge","Chinese LLMs not a monolith: aggressive play ranges 1-9%","Prisoner's dilemma: Chinese AI labs diverge more than East-West gap","Cooperation in Chinese AI varies by lab not country"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inference that the lab, not the model or ecosystem, is the unit of cooperative disposition assumes that each lab's single flagship model is representative of that lab's alignment lineage, since one model per lab at one point in time could instead reflect model-version or serving-backend effects.","fun_headline_variants_meta":{"raw":{"variants":["Chinese AI labs split on cooperation, new test shows","Lab, not ecosystem, sets AI cooperation: Chinese models diverge","Chinese LLMs not a monolith: aggressive play ranges 1-9%","Prisoner's dilemma: Chinese AI labs diverge more than East-West gap","Cooperation in Chinese AI varies by lab not country"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000851,"raw_usage":{"total_tokens":3816,"prompt_tokens":1176,"completion_tokens":2640,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":792,"completion_tokens_details":{"reasoning_tokens":2549}},"tokens_in":792,"tokens_out":2640,"duration_ms":18796,"temperature":1.0,"reasoning_tokens":2549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:11:31.217837+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the full fixed-converter protocol with two or more served models per Chinese lab under the same converter. If within-lab $P_A$ differences are as large as the between-lab spread, or if the lab ordering flips, the lab-level attribution fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the open all-play-all IPD tournament implementation used to run matches."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the finite-population selection model that generates the equilibrium proportions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Connects finite-population selection to cooperation in the IPD, motivating the equilibrium analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the action-noise mechanism used in the noisy conditions."},{"cited_title":"Hamilton","cited_arxiv_id":null,"evidence_quote":"Founds the IPD evolution-of-cooperation framework and its payoff structure."}],"review_version":1}