{"id":"7c6a0ebc-ed02-439e-8df5-76caba082e50","arxiv_id":"2607.14914","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In a stochastic ultimatum game with resource feedback, slow-growing resources sustain fairness through a spite-driven cycle: spite in depleted states replenishes the resource, after which repeated interactions favor fair divisions.","lead":"This paper models repeated bargaining between an owner and a worker over a self-renewing resource, and shows how slow resource recovery can create a cycle where spiteful rejections help replenish the resource, which then promotes fair offers. The result offers a possible explanation for why fairness and spite co-exist in natural systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing depleted-state payoff scale n leaves the central feedback loop unqualified: the cost of spite in the depleted state, and hence the claimed low-growth mechanism, cannot be evaluated or reproduced from the text.","rationale":"The paper is a self-contained modeling study with a clear emergent mechanism, and it provides code, which is genuine independent support. The reader's weakest assumption identifies n as an unparameterized factor in the depleted-state payoff matrix. I agree that this is the most load-bearing gap. The central claim is explicitly about a feedback loop in which spiteful behavior in the depleted state is favored because it transitions the resource to the replete state. Whether that tradeoff is evolutionarily favorable depends on the payoff difference between depleted and replete states, which is scaled by n. Without a specified n or a sensitivity analysis, the claimed mechanism is not fully pinned down: there will generally be a threshold n above which the future gains from replenishment no longer justify the immediate cost of spite, and below which the effect is artificially strong. This is a correctness risk rather than a disagreement with consensus: it concerns whether the reported phenomenon is robust or an artifact of an unspecified parameter choice. The concrete test—scanning n using the provided code—would settle whether the qualitative pattern in Figs. 3–4 persists. Since the reader's verdict is already CONDITIONAL and this concern is the same one, my stress-test pass does not move the verdict; it reinforces the need for the n-sensitivity analysis before full acceptance.","tokens_in":19389,"tokens_out":8937,"duration_ms":94383,"concrete_test":"Run the provided GitHub code (or an independent re-implementation) for τ0010 with all stated parameters (N_o=N_a=100, δ=0.99, w=0.5, h=0.5, l=0.05, μ_o=μ_a=0.01) and scan n over {0.05, 0.25, 0.5, 0.75, 0.95}. Record the stationary frequencies of the outcome profiles in Fig. 4 and the fairness/spite/replete levels in Fig. 3. If [FO_rep, SO_dep] remains a dominant profile and the fairness and spite levels change by less than ~10% across the scan, the concern is resolved. If, instead, the feedback profile loses dominance or the levels change sharply with n, the central claim must be restricted to the actual n used and accompanied by a threshold for n.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that for low growth rates (τ0010) a spite-driven feedback loop sustains fairness—rests on a payoff tradeoff: individuals must be willing to incur the cost of a spiteful (L,H) outcome in the depleted state to move the resource to the replete state, where larger payoffs make fairness worthwhile. The magnitude of that cost is controlled entirely by n, the depleted-state payoff scale introduced in §II and used in Eq. (A9): all depleted payoffs are multiples of n (e.g., n(1−h), n(1−l), nh, nl). Yet n is never assigned a numerical value anywhere in the text, and no sensitivity analysis is reported. If n is close to 1, depleted and replete payoffs are nearly equal, the future gain from replenishment may not compensate for the forgone payoff of agreeing in the depleted state, and the feedback loop could disappear or reverse. If n is close to 0, depleted interactions contribute almost nothing to fitness, so the selection pressure for the spiteful transition is artificially strong. Thus the headline result is conditional on an unstated value of n; Figs. 3–4 cannot be reproduced from the text alone, and the robustness of the loop to this parameter is unknown. The public code may contain a value, but the manuscript—the scientific record—does not specify it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a two-state stochastic mini-ultimatum game with repeated alternating interactions between an owner-offerer and an accepter. The resource switches between replete and depleted states according to one of two deterministic transition vectors, interpreted as high or low resource growth. The population consists of finite subpopulations of offerers and accepters using pure reactive strategies; a mutation-selection Markov chain built from simulated fixation probabilities yields long-run average fairness, spite, and resource-replete levels. The headline claim is that for low-growth resources (τ0010), spiteful (L,H) outcomes dominate in the depleted state, allowing the resource to replenish, and repeated interactions then promote fair (H,H) outcomes in the replete state, creating a self-sustaining spite-driven feedback loop; for high-growth resources (τ1111), fairness is maintained with little spite.","tokens_in":19814,"tokens_out":9668,"duration_ms":98211,"significance":"The manuscript extends the stochastic game-environment feedback literature, previously focused on cooperation, to the ultimatum game and to the joint evolution of spite and fairness. Strengths include an explicit Markov-chain derivation of fairness/spite rates for reactive strategies, a coherent two-species mutation-selection formulation, and publicly available code. If the numerical results are robust, the model offers a mechanistically plausible connection between resource scarcity, spite, and fairness. However, the quantitative results currently depend on an unstated depleted-state payoff parameter and on a qualitative rather than derived mapping from logistic growth to the transition vectors, so the significance is conditional.","major_comments":[{"comment":"The depleted-state payoff scaling parameter n (0<n<1) is introduced in Section II and appears in Eq. (A9), where all depleted-state payoffs are multiplied by n. The manuscript never assigns n a numerical value, and no sensitivity analysis is reported. This parameter controls the cost of the spiteful (L,H) action in the depleted state and therefore directly affects the selective pressure that underlies the claimed feedback loop. Without a value, Figs. 3-4 cannot be reproduced from the text, and the robustness of the loop to n is unknown. The authors should specify the value used in the code and report how the results vary with n, or derive n from the coarse-graining procedure.","section":"Section II; Eq. (A9)"},{"comment":"The identification of the transition vectors τ0010 and τ1111 with 'low' and 'high' resource growth rates is asserted through a heuristic interpretation rather than derived from the logistic equation. No mapping from r, m_thr, or the discretization time scale to the entries of τ is provided. Since the central claim is that resource growth rate controls whether the spite-driven feedback loop emerges, the paper currently demonstrates only that two hand-selected transition rules produce different outcomes. Please provide the coarse-graining calculation that yields these transition vectors, or at least an explicit parameter regime of the logistic model under which τ0010 and τ1111 arise.","section":"Section II.B"},{"comment":"The fixation probabilities used to build the mutation-selection Markov chain are estimated from birth-death simulations, but the number of realizations, random seed, and Monte Carlo uncertainty are not reported. Every long-run quantity in Figs. 3-4 inherits this estimation noise, so the reported levels carry unknown error bars. This is not necessarily fatal, but the manuscript should report the simulation effort and provide standard errors or confidence intervals for the reported fairness, spite, and replete levels.","section":"Appendix C; Figs. 3-4"}],"minor_comments":[{"comment":"The caption says the first and second columns correspond to τ0100 and τ1111, but the text uses τ0010 and τ1111. Please correct the typo.","section":"Fig. 3 caption"},{"comment":"Starting from (n_o,n_a)=(1,1), the two-species birth-death process also has the absorbing state (0,0), where both mutant types go extinct. The text states that there are exactly three absorbing states. Although the self-loop in the Markov chain can absorb the leftover probability, the statement is mathematically incorrect and should be revised.","section":"Appendix C"},{"comment":"'To crystal clearly illustrate' should read 'To crisply illustrate' or similar.","section":"Section II.B"},{"comment":"'a spite and fairness driven resource feedback loop' should be hyphenated for clarity ('spite-and-fairness-driven').","section":"Section V"},{"comment":"Color maps are used without a color bar. Since the numerical values of the levels are central, adding a color scale or numeric labels would improve reproducibility and readability.","section":"Figs. 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"The missing parameter n is likely present in the public repository, but the manuscript itself must specify it and demonstrate robustness. The paper is a plausible extension of the authors' earlier mini-UG framework, and the central idea is worth publishing if the sensitivity and derivation details are supplied. The fit with nlin.AO is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. The paper delivers a genuinely new result: a stochastic ultimatum game with a coarse-grained renewable resource, where a spite-driven feedback loop sustains fairness for slow-growing resources. That is not just a relabeling of existing stochastic-game work, which has been mostly about Prisoner's Dilemma. The model is coherent, the Markov-chain derivations in the appendices check out, and the authors are careful about the mutation–selection setup. Public code is a real plus.\n\nThe soft spot is the one the stress-test flags, and it is real. The parameter n—the depleted-state payoff scale, 0<n<1—multiplies every payoff in the depleted state (Eq. A9). It controls exactly how costly being spiteful is when the resource is down: a low n makes depleted interactions nearly irrelevant to fitness, so selection to move out of that state is artificially strong; an n close to 1 makes depleted and replete payoffs comparable, so the future gain from replenishment may not compensate the cost of spite, and the loop could vanish. The manuscript defines n but never assigns it a value and reports no sensitivity analysis. Without that, the headline result is conditional on an unstated parameter, and Figs. 3–4 cannot be reproduced from the text alone. This is a fixable omission, not a structural flaw. I checked the GitHub link in spirit; the code may well contain a value, but the scientific record doesn't.\n\nA secondary weakness is the mapping from the continuous logistic resource to the two deterministic transition vectors. The threshold m_thr is not specified, and the reduction is heuristic. That is a modeling choice, not an error, but it would be good to see a discussion of how the coarse-graining affects the results.\n\nThe paper has no data fitting, so there is no circularity. The self-citations are to the authors' own repeated mini-UG framework, which is a legitimate building block. The fixation probabilities are simulated, and no error bars are reported; that is a minor point.\n\nVerdict: worth refereeing. The idea is new, the framework is transportable, and the missing n is exactly what a competent referee would ask for. I would recommend sending it out, with the expectation of moderate revision—specify n or run it across a range, add error bars or at least state the number of realizations, and address the coarse-graining. I'd bring it to a reading group, and if I worked on evolution of fairness or stochastic games, I'd cite the mechanism once the parameter issue is settled.","headline":"Genuinely new mechanism for spite–fairness coupling via resource feedback, but the missing value of the depleted-state payoff scale n makes the headline result conditional.","tokens_in":20205,"tokens_out":3365,"would_cite":true,"duration_ms":32848,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A22","92D15","60J20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Slow resource growth turns spite into the engine that restores abundance and selects for fairness in a repeated ultimatum game.","keywords":["stochastic ultimatum game","resource feedback","spite","fairness","reactive strategies","mutation-selection dynamics","game-environment feedback","two-species population dynamics"],"falsifier":"Vary n over (0,1) in the fixation-probability calculations and re-compute the long-run mutation–selection distribution; if the spite-fairness loop vanishes for some n, the result is an artifact of an unstated parameter choice. Alternatively, in a laboratory repeated ultimatum game with a slowly replenishing endowment, check whether fair offers and spiteful rejections actually alternate with endowment abundance as the model predicts.","tokens_in":19344,"feed_emoji":"♻️","tokens_out":7916,"duration_ms":67321,"temperature":0.7,"pith_summary":"This paper sets out to establish that spite and fairness can be two phases of a single resource-driven evolutionary cycle rather than mutually exclusive outcomes. The model is a repeated ultimatum game between an owner-offerer and an accepter, with a self-renewing resource that is either replete or depleted; the resource state rescales the payoffs, and the players' agreements or rejections in turn change the resource state. For resources with high growth rates, fairness dominates and spite stays low. For slowly growing resources, the long-run outcome is a spite-driven feedback loop: spiteful rejections dominate in the depleted state, allow the resource to recover, and the resulting abundance together with repeated interaction selects for fairness, which then depletes the resource again. The paper's contribution is a mechanism—built only from game-environment feedback and mutation-selection dynamics—that explains how costly antisocial behaviour can persist and even sustain prosocial fairness.","feed_headline":"Spite drives fairness when resources grow slowly","feed_subtitle":"When a shared resource regrows slowly, spiteful rejections restore it, and repeated play then favors fair offers.","key_machinery":"The load-bearing object is the two-state stochastic mini-ultimatum game with deterministic transition vectors τ0010 and τ1111. These specify how the resource responds to each action pair: τ0010 represents a slowly growing resource, where successful agreements keep the resource depleted and only the spiteful (L,H) pair lets it recover; τ1111 represents a fast-growing resource, where any play in the depleted state restores it to replete. The argument runs through two nested Markov chains—an 8-state chain over resource-plus-action states that yields average payoffs, spite rates, and fairness rates under discounting, and a 256-state mutation–selection chain over strategy pairs that yields the lo","core_discovery":"The central claim is that a slowly renewing resource converts the ultimatum game into a self-sustaining alternation between two behavioural regimes. In the depleted state the dominant outcome is the spiteful action pair (low offer, high demand), which blocks agreement and therefore halts exploitation; since the resource is self-renewing, the pause lets it return to the replete state. In the replete state, repeated interactions make the fair outcome (high offer, high demand) dominant—offering high avoids rejection, and demanding high builds a reputation that deters low offers. Fair agreements then harvest the resource and drive it back to depletion, closing the loop. The paper establishes thi","pith_inferences":["If the loop is robust, managing a resource so that it stays moderately depleted could, in the short term, select for spiteful conflict before fairness returns; this is an inference beyond what the paper asserts.","A direct check the paper leaves open: the depleted-state payoff scaling factor n is never given a value, so re-running the Markov chain across n in (0,1) would test whether the loop survives all choices.","The owner–non-owner asymmetry may itself stabilise the cycle; modelling symmetric roles could reveal whether ownership is essential to the feedback or just a convenient frame.","A laboratory test with a slowly replenishing stake should show alternating fair offers and spiteful rejections tied to stake abundance, rather than a fixed bargaining norm."],"forward_implications":["For slowly growing resources, the evolutionary outcome is not a single strategy but a cycle: depleted states select for spite, replete states select for fairness, so both behaviours coexist across time.","The model predicts that in a depleted environment, costly spiteful rejections function as a repair step, restoring the resource and creating the abundance under which fairness thrives.","The framework carries game–environment feedback beyond cooperation in the Prisoner's Dilemma into the ultimatum game, making spite and fairness themselves environmentally modulated traits.","The difference between the two transition vectors implies that the resource growth rate, not an intrinsic taste for fairness, can decide whether a population looks fair or oscillates between spite and fairness.","The results put qualitative experimental observations—scarcity triggering antisocial behaviour and repeated interactions promoting fairness—into one quantitative model."],"fun_headline_variants":["Slow regrowth turns spite into fairness","Spiteful rejections restore resources, then fairness wins","Sparse resources use spite to force fair deals","When resources lag, spite and fairness alternate"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The loop rests on the depleted-state payoffs being a fixed fraction n (0<n<1) of the replete-state payoffs, but the paper never specifies n; because all selection in the depleted state flows through those payoff differences, the reported patterns in Figs. 3–4 cannot be reproduced from the text and may change for other n.","fun_headline_variants_meta":{"raw":{"variants":["Slow regrowth turns spite into fairness","Spiteful rejections restore resources, then fairness wins","Sparse resources use spite to force fair deals","When resources lag, spite and fairness alternate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1204,"prompt_tokens":707,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":439}},"tokens_in":451,"tokens_out":497,"duration_ms":5297,"temperature":1.0,"reasoning_tokens":439,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:41:15.087244+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Vary n over (0,1) in the fixation-probability calculations and re-compute the long-run mutation–selection distribution; if the spite-fairness loop vanishes for some n, the result is an artifact of an unstated parameter choice. Alternatively, in a laboratory repeated ultimatum game with a slowly replenishing endowment, check whether fair offers and spiteful rejections actually alternate with endowment abundance as the model predicts.","supporting_citations":[],"review_version":1}