{"id":"ac0f18f3-f2b5-41f4-aaad-1c00ce6782d6","arxiv_id":"2508.08301","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In the coordination-II payoff ordering (1 > 0 > S > T), a population of one-shot sampling players (PR(1)) is an evolutionarily stable strategy that resists invasion by fully von Neumann-Morgenstern rational mutants, making bounded rationality the evolutionarily stable basis of cooperation.","lead":"This paper argues that in a modified stag-hunt game where cheating is especially frowned upon, players using a simple one-sample decision rule are evolutionarily stable against more rational players, so cooperation can persist because evolution favors less rational individuals. A reader might care because it connects behavioral economics with evolutionary game theory and suggests bounded rationality is an evolutionary advantage.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mean-field closure in Eq. (2) is unvalidated for PR(1)-VR and VR-VR pairs; if it mispredicts asymptotic beliefs, the payoff matrix Π and PR(1)'s ESS status could be wrong.","rationale":"The reader identified the same concern, and I agree it is the weakest spot. The paper is transparent about the approximation, but the validation is incomplete; the closure for the VR best-response is particularly risky because the true expectation of a threshold response depends on the full distribution of beliefs, not just the mean. The numerical Fig. 1 does not resolve this, as it uses the same mean-field assumption. No other issue is as directly tied to the central claim. The code is available and the approximation is honestly disclosed, but conditional acceptance with a request for cross-type validation is appropriate; hence the verdict should remain CONDITIONAL (UNCHANGED).","tokens_in":14278,"tokens_out":6660,"duration_ms":73474,"concrete_test":"Simulate the exact stochastic belief-update rule (Eq. 1) for PR(1)-VR and VR-VR pairs on a fine grid of initial beliefs (x0,y0) ∈ [0,1]^2, at multiple points across the coordination-II parameter space (e.g., S,T sampled from -1 < S < 0, -1 < T < S). For each pair type and parameter point, compute the empirical distribution of asymptotic beliefs and the resulting average payoff over initial beliefs, producing a stochastic payoff matrix Π_stoch. Then check the ESS conditions for PR(1) against VR using Π_stoch. If PR(1) is ESS at all tested points, the mean-field closure is adequate for the central claim; if not, the paper's Fig. 1 is unreliable and the claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that PR(1) is an ESS in coordination-II games (Fig. 1) is computed from the payoff matrix Π of Section IV, whose entries are asymptotic beliefs from the mean-field dynamics Eq. (2). Eq. (2) replaces ⟨BR(y)⟩ with BR(⟨y⟩). For PR(1), BR is smooth (2y - y^2), but for VR, BR is a step function (1 above the mixed-Nash threshold, 0 below, 1/2 exactly at threshold). Averaging a step function over the distribution of y is not generally the step function of the average; if beliefs fluctuate across the threshold, the mean-field VR response is biased. Appendix A validates the mean-field closure only for two PR(1) players, and only at one parameter point (S = -0.5, T = -1.5), showing the qualitative outcome (convergence to (1,1)) for that pair. No simulation is given for PR(1)-VR or VR-VR pairs that contribute to Π. If the mean-field ODE mispredicts whether these mixed pairs converge to (0,0) or (1,1) for a substantial fraction of initial beliefs, then the averaged payoffs Π(PR,VR), Π(VR,PR), and Π(VR,VR) are wrong, and the ESS inequalities (Π(PR,PR) > Π(VR,PR), or equality plus Π(PR,VR) > Π(VR,VR)) could flip. Since the ESS claim is the paper's central result, this unresolved approximation is the most load-bearing concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the evolution of cooperation in a coordination-II modification of the stag-hunt game, defined by the payoff ordering 1>0>S>T. It contrasts VNM-rational (VR) players with procedurally rational PR(k) players, who estimate action utilities by sampling each action k times. The authors propose a belief-updating dynamics (Eq. 1), pass to a mean-field ODE (Eq. 2), and compute, by numerical averaging over initial beliefs, the payoff matrix Π for PR(1) versus VR. They then apply standard ESS criteria and report in Fig. 1 that PR(1) is an ESS over the entire coordination-II parameter region, while VR is not except in a small region. Appendices extend the claim to finite populations (ESS_N), stochastic stability, PR(2) mutants, and alternative belief-update rules. The paper concludes that bounded rationality can evolutionarily stabilize cooperation.","tokens_in":14693,"tokens_out":17289,"duration_ms":189949,"significance":"If the central claim holds, the paper is significant and novel: it offers a mechanism by which less rational decision-making is evolutionarily advantageous and provides a behavioral-game-theoretic resolution of the stag-hunt dilemma. The model is explicit, the results are numerically checkable, and the authors make the code available on GitHub; the ESS region is a concrete falsifiable prediction. However, the main quantitative conclusion currently rests on an unvalidated mean-field closure for mixed rationality types, and the robustness appendices contain gaps. These issues should be addressed before the claims can be fully accepted.","major_comments":[{"comment":"The payoff matrix Π — and hence the central ESS claim — is obtained from the mean-field equations (2). The closure ⟨BR(y)⟩ ≈ BR(⟨y⟩) is validated in Appendix A only for two PR(1) players at S=-0.5, T=-1.5 (Fig. 2). No stochastic simulation of Eq. (1) is shown for the PR(1)-VR and VR-PR pairs that determine Π(PR,VR) and Π(VR,PR), nor for PR(2)-involving pairs used in Fig. 1(a). For VR, BR is a step function, so fluctuations of the opponent's belief across the mixed-NE threshold can make E[BR(y)] very different from BR(E[y]); the averaged payoffs and the basin structure in Fig. 1(b) may therefore be wrong. Please add stochastic validation for the mixed-type interactions across the parameter range, or state clearly that Fig. 1 is a mean-field prediction and reassess the ESS status if the closure fails.","section":"Section IV and Appendix A"},{"comment":"The proof that PR is ESS_N for any population size states: 'since A corresponds to coordination-II game, we have a > d, b = c and a > b'. For the underlying payoff matrix A = [[1,S],[T,0]], b=c means S=T, which is not true in the coordination-II region (1>0>S>T, S≠T generally). The type-level payoff matrix Π of Section IV is not symmetric either. Unless b=c is proved from the averaging over initial beliefs, the two ESS_N inequalities do not follow. Please correct the argument or replace the finite-population claim with a numerical check.","section":"Appendix C"},{"comment":"The stochastic model in Eq. (D1) first imposes Γ(0)=Γ(1)=0 to ensure forward invariance of p∈[0,1], but then assumes Γ(p)=σ constant. These conditions are incompatible unless σ=0. The Fokker-Planck solution (D2) and the conclusion that p=1 is the stochastically stable equilibrium rely on the noise model. Please clarify the boundary treatment (e.g., use reflecting boundaries and admit Γ constant, or make Γ vanish at the boundaries) and re-derive the SSE claim.","section":"Appendix D"}],"minor_comments":[{"comment":"The transition from the stochastic difference equation (1) to the differential equation for the mean is not fully explained. In particular, the notation BR(⟨y⟩) for PR(1) should be explicitly defined as the expected best response (i.e., the probability of choosing C), since in Eq. (1) BR is a random action. The current notation conflates the random variable and its expectation.","section":"Eq. (2)"},{"comment":"The word 'unequivocally' overstates what is shown: Fig. 1 is a numerical scan over the S-T plane without reported error bars or resolution analysis. Consider softening to 'numerically' or 'within the model's mean-field approximation'.","section":"Abstract and Section IV"},{"comment":"The caption and text refer to 'dashed line' and 'dotted line', but the figure is not reproduced in the manuscript text. Ensure the lines are clearly labelled in the final figure.","section":"Figure 1"},{"comment":"The payoff matrix notation ⟨yPR∞⟩·A⟨xPR∞⟩ is hard to parse. Define x∞ and y∞ for each type combination explicitly (e.g., in a small table) before presenting the matrix.","section":"Section IV"},{"comment":"The statement 'Our conclusion remains unchanged for any other such game' is a claim, not a demonstration, since the validation is shown for one parameter point. Please state the range of S,T checked or remove the generalization.","section":"Appendix A"},{"comment":"Reference [57] (Mukhopadhyay & Chakraborty, Chaos 2021) is cited for the replicator equation but appears unrelated to that specific statement; citing a standard textbook or the original replicator equation papers [53,56], already present, would be more appropriate.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The mean-field validation gap is the main risk: the central ESS conclusion depends on Π entries for PR(1)-VR pairs that are not validated against the stochastic dynamics. If the authors can supply those simulations and the ESS region is reproduced, the paper is likely publishable. The Appendix C and D errors are fixable but need correction. Also note that reference [57] appears to be an unrelated self-citation; this may merit editorial attention."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a genuinely new idea: treat rationality level as an evolvable trait and ask whether a population of one-sample procedural rationalists (PR(1)) can resist invasion by more rational types in a coordination-II game. The setup is explicit, the averaging over initial beliefs is clearly specified, and PR(1) being an ESS across the coordination-II parameter space is not in the cited literature. The code is available, and there are extra robustness checks (finite populations, stochastic stability, random-belief mutants). That is real work and deserves credit.\n\nThe soft spot is the one the report flags. The mean-field closure in Eq. (2) replaces <BR(y)> with BR(<y>). For PR(1), BR is smooth, so the closure is tolerable. For VR, BR is a step function at the mixed-Nash threshold. If beliefs fluctuate near that threshold, the average of the step can differ from the step of the average. Appendix A validates the closure only for two PR(1) players. The payoff matrix Pi, which drives the ESS result, includes PR(1)-VR and VR-VR pairs where the VR step matters, and no stochastic simulation is shown for those pairs. If the mean-field mispredicts the fraction of initial beliefs that end at (1,1) vs (0,0), the payoff entries shift and the ESS phase diagram could change. This is not a minor detail; it is load-bearing for the central claim.\n\nA few smaller things. The abstract says \"unequivocally,\" which is too strong for a numerical result. The payoff matrix table is hard to read and appears to have off-diagonal entries swapped; if that is a typo, it should be fixed. Appendix C asserts b=c for the fitness matrix, but for an asymmetric underlying payoff (S != T) the off-diagonal payoffs are not obviously equal; that needs justification.\n\nThe central idea is good and the model is carefully specified. The paper is not incoherent, but the main result is only as solid as the unvalidated approximation. I would send it to peer review with a request for the authors to simulate the mixed-type interactions directly and report error bars or an analytical argument for the closure. If that works out, it is a nice contribution. If not, the ESS claim may not survive.\n\nRecommendation: accept for review, but the referees should push on the mean-field.","headline":"New idea, interesting result, but the key claim rests on an unvalidated mean-field approximation for mixed rationality pairs; with that fixed, it's a solid paper.","tokens_in":15098,"tokens_out":6645,"would_cite":false,"duration_ms":73807,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A22","91A05"],"pacs":[],"model":"deepseek-v4-flash","headline":"In the stag-hunt dilemma, evolutionary stability belongs to bounded rationality, not full rationality: a population of one-shot samplers (procedurally rational players) resists invasion by fully rational mutants throughout the coordination-","keywords":["coordination games","stag-hunt game","bounded rationality","procedural rationality","evolutionarily stable strategy","cooperation","sampling equilibrium","evolutionary game theory"],"falsifier":"Run the fully stochastic belief-updating rules, Eqs. (1), for PR(1)-VR and VR-VR pairs across a grid of $(x_0, y_0)$ and $(S, T)$ in the coordination-II region, compute the resulting asymptotic beliefs and the payoff matrix $\\Pi$ without the mean-field approximation, and test the ESS inequalities; if any region of the $S$–$T$ plane admits a VR mutant that satisfies the invasion conditions against the simulated PR(1) resident, the paper's central claim fails as stated.","tokens_in":14205,"feed_emoji":"🦌","tokens_out":9471,"duration_ms":97354,"temperature":0.7,"pith_summary":"This paper argues that cooperation in hunter-gatherer-style stag-hunt dilemmas may have been secured not despite, but because of, human bounded rationality. In the coordination-II game — a stag-hunt variant with payoff ordering $1 > 0 > S > T$, reflecting a preference for being cheated over cheating — the authors show that a population of procedurally rational players, who mentally sample each action once before choosing (PR(1)), is an evolutionarily stable strategy. Such a population resists invasion by more rational mutants, including fully rational expected-utility maximizers, across the whole coordination-II parameter range, and the stability survives finite population size, continuous mutation, and mutants with random belief updates. The upshot is a reversal of the usual default: full rationality is the fragile trait, and bounded rationality is what evolutionary forces select.","feed_headline":"Bounded rationality defends cooperation in the stag-hunt game","feed_subtitle":"One-shot samplers resist invasion by fully rational mutants, so evolution picks bounded rationality.","key_machinery":"The load-bearing objects are the belief-update recursions $x_{m+1} = \\frac{m}{m+1} x_m + \\frac{1}{m+1} BR(y_m)$ (and symmetrically for $y$), which formalize fictitious-play-style virtual experimentation for procedural samplers; the mean-field approximation $\\langle BR(y)\\rangle \\approx BR(\\langle y\\rangle)$ that closes Eq. (2) and turns the recursions into differential equations; and the identification of PR(∞) with VR, which places full rationality at the infinite-sampling endpoint of the PR(k) family. The fitness matrix $\\Pi$, with each entry the double average (over stochasticity and over initial beliefs) of asymptotic payoffs, is what the ESS inequalities are evaluated on.","core_discovery":"The central claim is that PR(1) — the player who samples each action exactly once and picks the better sampled outcome — is an evolutionarily stable strategy (ESS) against PR(2) and PR(∞) = VR mutants over the whole coordination-II parameter range $1 > 0 > S > T$. The case rests on the belief-update dynamics of Eq. (1), the mean-field closure $\\langle BR(y)\\rangle \\approx BR(\\langle y\\rangle)$, and the fitness matrix $\\Pi$ averaged over initial beliefs. The key asymmetry: two PR(1) players reach the cooperative equilibrium $(1,1)$ from almost every initial belief, while VR pairs end at $(1,1)$ or the defection equilibrium $(0,0)$ depending on initial conditions — so PR(1) outscores VR in the","pith_inferences":["(Editorial inference) The mean-field closure $\\langle BR(y)\\rangle \\approx BR(\\langle y\\rangle)$ is validated in the paper only for two PR(1) players; a direct stochastic simulation of Eqs. (1) for PR(1)-VR and VR-VR pairs, without the closure, is the natural check that the reported ESS boundary in the S-T plane is not an artifact of that approximation.","(Editorial inference) If sampling cost is treated as a fitness penalty, PR(1) should become only more dominant, since it reaches the same cooperative payoff as PR(k>1) and VR while spending the least cognitive effort; adding a cost term to $\\Pi$ is a direct extension.","(Editorial inference) The paper's logic suggests a behavioural prediction: in laboratory coordination-II games, players should systematically converge to the payoff-dominant cooperative equilibrium, and treatments that nudge subjects toward expected-utility reasoning should show more defection — a pattern consistent with the sampling-equilibrium experiments the paper cites.","(Editorial inference) The PR(k) family embeds the ESS question in a one-parameter ladder between bounded and full rationality; a continuous-strategy version (real-valued sampling effort) would let one ask whether the ESS result is robust to arbitrarily small increases in rationality."],"forward_implications":["In the coordination-II game, a resident PR(1) population cannot be invaded by any mutant that samples more times per action, including the fully rational VR mutant; cooperation stays in place.","A fully rational population is not safe: for most of the S-T parameter space VR is not ESS, so procedurally rational mutants can invade and restore cooperation.","The stability of PR(1) is not an artifact of the infinite-population idealization: it survives in finite populations of any size (ESS$_N$) under the Moran process.","PR(1) is also the unique stochastically stable equilibrium against continuously appearing mutants, while VR is never stochastically stable — noise in the mutation process cannot dislodge bounded rationality.","Because different PR(k) correspond to different sampling effort, the result ties the evolution of rationality to cognitive cost: the level of rationality that survives is the minimal one that still reaches the cooperative outcome."],"supporting_citations":[{"why":"Provides the experimental evidence that sampling equilibrium describes human behaviour better than Nash equilibrium, motivating the procedural-rationality model.","marker":"[45]"},{"why":"Supplies the sampling dynamics whose SE(1) outcome the PR(1) belief-update rule is built on, including the result that only the cooperative equilibrium is reached in coordination-II games.","marker":"[48]"},{"why":"Names and defines the coordination-II game, the payoff structure ($1 > 0 > S > T$) on which all the paper's results rest.","marker":"[40]"},{"why":"Gives the ESS/replicator-dynamics framework that connects evolutionary stability to asymptotic stability of the replicator equation, justifying the payoff-matrix ESS criterion.","marker":"[53]"},{"why":"Provides the ESS_N conditions used to extend PR(1) stability to finite populations of any size under the Moran process.","marker":"[59]"},{"why":"Supplies the stochastic-stability machinery used to show PR(1) survives continuous mutation while VR does not.","marker":"[60]"}],"fun_headline_variants":["Procedural rationality emerges as ESS in stag-hunt","Bounded rationality stabilizes cooperation in stag-hunt","One-shot samplers resist rational mutants in stag-hunt","Evolution picks bounded rationality for coordination","Cooperation emerges from bounded rationality in stag-hunt"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The central result assumes that a player's average best response equals her best response to her average belief — a closure checked for two one-shot samplers but assumed without simulation for interactions between different rationality types; if it fails there, the fitness matrix and the claimed stability could change.","fun_headline_variants_meta":{"raw":{"variants":["Procedural rationality emerges as ESS in stag-hunt","Bounded rationality stabilizes cooperation in stag-hunt","One-shot samplers resist rational mutants in stag-hunt","Evolution picks bounded rationality for coordination","Cooperation emerges from bounded rationality in stag-hunt"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1500,"prompt_tokens":688,"completion_tokens":812,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":737}},"tokens_in":432,"tokens_out":812,"duration_ms":7583,"temperature":1.0,"reasoning_tokens":737,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:35:19.775023+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the fully stochastic belief-updating rules, Eqs. (1), for PR(1)-VR and VR-VR pairs across a grid of $(x_0, y_0)$ and $(S, T)$ in the coordination-II region, compute the resulting asymptotic beliefs and the payoff matrix $\\Pi$ without the mean-field approximation, and test the ESS inequalities; if any region of the $S$–$T$ plane admits a VR mutant that satisfies the invasion conditions against the simulated PR(1) resident, the paper's central claim fails as stated.","supporting_citations":[{"cited_title":"Cosmides, Cognition31, 187–276 (1989)","cited_arxiv_id":null,"evidence_quote":"Provides the experimental evidence that sampling equilibrium describes human behaviour better than Nash equilibrium, motivating the procedural-rationality model."},{"cited_title":"Hummert, K","cited_arxiv_id":null,"evidence_quote":"Supplies the sampling dynamics whose SE(1) outcome the PR(1) belief-update rule is built on, including the result that only the cooperative equilibrium is reached in coordination-II games."},{"cited_title":"Hill, Human Nature13, 105–128 (2002)","cited_arxiv_id":null,"evidence_quote":"Names and defines the coordination-II game, the payoff structure ($1 > 0 > S > T$) on which all the paper's results rest."},{"cited_title":"Selten and T","cited_arxiv_id":null,"evidence_quote":"Gives the ESS/replicator-dynamics framework that connects evolutionary stability to asymptotic stability of the replicator equation, justifying the payoff-matrix ESS criterion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ESS_N conditions used to extend PR(1) stability to finite populations of any size under the Moran process."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the stochastic-stability machinery used to show PR(1) survives continuous mutation while VR does not."}],"review_version":1}