{"id":"e5865706-6c3d-4c16-8020-8f01a9dc2ed1","arxiv_id":"2508.20799","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In an evolutionary model of the Incremental Centipede Game, positively biased level-1 reasoning outcompetes rational, unbiased, and negatively biased strategies, and can coexist with myopic no-reasoning players.","lead":"This paper models how cognitive biases and reasoning depth co-evolve in a population playing the Incremental Centipede Game. It finds that evolution consistently favors positively biased reasoning, while rational play goes extinct, and longer games amplify the effect.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bias is implemented as directional action noise, not inference bias; B+ dominance may depend on the ad hoc kernel shape in Appendix B.","rationale":"The paper is internally consistent and the Appendix B derivation of σ(k)=σ(0)M(ε)^k is sound; the code is openly available and the results are reproducible in principle. The reader correctly identified the construction of the biased kernels as the weakest assumption. My stress-test sharpens this: the bias is implemented as directional action noise, not as an inference or belief bias, so the abstract's language about 'systematic inference bias' and 'overestimating the probability of later termination' is not directly instantiated. Because the main result—dominance of B+(1) and coexistence with NR—could plausibly change under alternative kernel implementations, and because no robustness analysis is supplied, the central claim should remain conditional. The empirical fitting concern is secondary: even if the fit is in-sample, the evolutionary result stands as a model result over the parameter ranges shown; it is the kernel-shape dependence that threatens the interpretation of the central claim. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":20166,"tokens_out":10054,"duration_ms":105221,"concrete_test":"Re-run the Fig. 2 stationary-distribution analysis (L=6, β∈[10^-4,10], ε∈[0.05,0.3]) with B+(k) redefined as a belief-based bias: the agent uses the unbiased action kernel M_U but treats the opponent as level k-2 instead of k-1, while B-(k) continues to assume level k. If B+(1) is no longer the dominant strategy under strong selection and the NR/B+(1) coexistence point in Fig. 3 disappears, the headline result is an artifact of the Appendix B action-tremble construction rather than a property of positively biased inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is the construction of the biased kernels in Appendix B. The paper's headline claim is that selection favours a *reasoning* bias (wishful thinking, overestimation of later termination). But the implemented B+/B- kernels are not belief biases: they are directional action trembles. In M_B+, the ε deviation mass is placed only on later nodes; in M_B-, only on earlier nodes, uniformly over those nodes. The recursive inference step (Eqs. 2 and 3) is identical across U, B+, and B-: the agent always assumes the opponent is level k-1. Thus the model demonstrates that an action bias toward later stopping is favoured, not that an inference bias is favoured. This distinction matters because the B+(1) advantage and the NR/B+(1) coexistence in Fig. 3 may follow directly from the asymmetric placement of the noise mass: B+ pushes play toward high-payoff later nodes whenever an earlier best response exists, while B- is constrained to move into early, low-payoff nodes. No sensitivity analysis is provided against alternative, equally natural implementations, such as biasing the belief about the opponent's level or using non-uniform deviation probabilities. The paper's own Conclusions acknowledge the static-belief framework and the 'unequivocally translated' assumption, but do not test whether the result survives a change in kernel shape. Since the central claim is phrased as a general evolutionary explanation of cognitive bias, the ad hoc kernel shape is the least secure load-bearing element.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the co-evolution of reasoning depth and cognitive biases in the Incremental Centipede Game. It introduces an unbiased level-k kernel and two biased kernels (B+ and B-), defined by a noisy best response where deviations to later (or earlier) nodes are allowed, and computes exact payoffs and evolutionary dynamics under the small-mutation limit. Results: under strong selection, B+ becomes dominant, reasoning depth remains bounded, SPE goes extinct, B+(1) and no-reasoning (NR) can coexist, and longer games enlarge the parameter region favoring B+. The model is calibrated to two experimental datasets.","tokens_in":20526,"tokens_out":6936,"duration_ms":72292,"significance":"If the result holds, the paper offers an evolutionary explanation for systematic deviations from rationality in a canonical sequential game without invoking other-regarding preferences. Strengths include a fully specified model, analytic derivation of sigma(k)=sigma(0)M^k, exact stationary distributions, open-source code, and calibration to two experimental datasets. The main weakness is that the bias is implemented as directional action noise rather than as an inference bias, so the headline conceptual interpretation is not directly supported by the model.","major_comments":[{"comment":"The central claim of a 'systematic inference bias' is not what the model implements. In Eq. (2), all kernels use the same recursive inference (opponent is level k−1); the kernels differ only in which actions the ε-noise can reach. The Appendix B matrices show that MB+ moves probability only to later nodes and MB− only to earlier nodes. Thus the model shows that a directional action bias toward later, higher-payoff nodes is selected, not that agents overestimate the probability of later termination. The Conclusions acknowledge 'unequivocally translated' behavior and static beliefs, but no test of an actual belief bias is provided. Please add variants (e.g., biased prior over opponent's level, non-uniform deviation probabilities) and report whether the qualitative results survive; alternatively, reframe the claims as about action bias.","section":"Methods §4.2, Eq. (2)–(3); Appendix B; Conclusions"},{"comment":"The coexistence equilibrium is identified in a replicator equation restricted to three strategies (B+(1), NR, U(1)). The full strategy set includes 20 strategies for L=6, and Panel A shows no ERS and a cycle. The 'new stable equilibrium' claim needs either an invasion analysis against all remaining strategies or an explicit statement that the coexistence point is an equilibrium of the reduced 3-strategy subsystem. The full Markov chain with mutations (Panel D) is a mutation–selection balance, not a deterministic attractor; please clarify the status.","section":"Results §2, Figure 3B"}],"minor_comments":[{"comment":"The phrase 'inference bias' appears in the Abstract and Introduction, but the implemented bias is an action tremble; align terminology throughout.","section":"Title/Abstract"},{"comment":"Several OCR-like typos in figure text ('no reaso ing', 'subgame perfect eq.') and garbled axis labels in Fig. 5 should be corrected.","section":"Figure 2 caption"},{"comment":"Please define concisely what happens at indifference: Appendix B sets the first row of each M matrix to uniform because Player 2 has no best response when T1=0; this should be explained in the main text.","section":"Methods §4.2"},{"comment":"Panel D of Fig. 3 uses μ=0.01 while the rest of the paper uses the small-mutation limit; clarify why this value is chosen and whether results are robust to μ.","section":"Methods §4.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript header states it is an early version of a paper published in Royal Society Interface (DOI provided). If this is a fresh submission, the editor should clarify prior-publication policy and dual-submission concerns. The paper's contribution is solid as a modeling exercise, but the interpretation issue in the report is central."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know before reading: this is a properly specified, reproducible evolutionary model, and it does show that a particular kind of directional action noise — always trembling toward later centipede nodes — is selectively favored. But the headline claim that this is \"positively biased reasoning\" in the sense of wishful thinking overshoots. The bias sits in the action kernel, not in the recursive beliefs: agents with B+(k) and U(k) all assume the opponent is level k-1; the only difference is where the epsilon mass is placed. So the model gives you a mechanism for biased action, not a mechanism for biased inference.\n\nWhat's genuinely new: the three reasoning kernels and the co-evolution of kernel and depth, which goes beyond the authors' earlier shared-reasoning model and Rand & Nowak's pure-strategy analysis. The coexistence equilibrium between B+(1) and NR is a new dynamical object. The math in Appendix B is coherent, the stationary distributions are derived cleanly, and there's working code. The qualitative prediction that longer games widen the window of epsilon where the positive kernel dominates is a real, falsifiable model result.\n\nSoft spots, in order. First, the kernel construction in Appendix B is ad hoc. Deviation mass is placed only on later nodes (B+) or earlier nodes (B-), spread uniformly. The stress-test note is right: this is an action bias, not an inference bias, and the main results may depend on this exact shape. The paper's own Conclusions acknowledge the \"unequivocally translated\" assumption and the static-belief framework, but there's no sensitivity analysis. A check with a different deviation distribution or a belief-level bias would tell you whether the qualitative result is robust. My hunch is the direction of the bias matters more than the uniform spread, but that's not demonstrated.\n\nSecond, the empirical \"agreement\" with the two experimental datasets is a calibration, not a prediction. The parameters beta and epsilon are fit to those datasets, so the resulting match should not be sold as independent support. The cross-cultural difference in fitted epsilon is interesting, but it is not a strong test.\n\nNeither issue invalidates the central result. The model is internally consistent and the selection result is real. The interpretation is just bigger than the model supports.\n\nAudience: evolutionary game theorists and anyone working on level-k models of bounded rationality. It is a solid contribution, not a game-changer. Worth a serious referee, provided the authors add a robustness analysis and trim the cognitive interpretation.","headline":"A clean, reproducible evolutionary model of directional action noise that is being oversold as a mechanism for cognitive bias; the result deserves a serious referee but needs a robustness analysis and a toned-down interpretation.","tokens_in":21004,"tokens_out":3650,"would_cite":false,"duration_ms":38409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A22","91A26","91A18"],"pacs":[],"model":"deepseek-v4-flash","headline":"Reasoning that is biased toward larger future rewards outcompetes rational backward induction in the centipede game, driving full rationality to extinction in the model.","keywords":["evolutionary game theory","centipede game","level-k reasoning","cognitive bias","positivity bias","bounded rationality","wishful thinking","replicator dynamics"],"falsifier":"Run the same evolutionary dynamics with the positively biased kernel replaced by a kernel that biases beliefs about the opponent's reasoning level instead of biasing the action deviation; if positive bias no longer dominates, the result rests on the specific kernel shape. Behaviourally, the model predicts the noise window for majority positive bias widens with game length, so comparing stopping-node distributions across Incremental Centipede Games of lengths 4, 6, and 8 with identical per-step stakes would test the mechanism.","tokens_in":20031,"feed_emoji":"🧠","tokens_out":5735,"duration_ms":62677,"temperature":0.7,"pith_summary":"This paper asks whether cognitive biases that make people expect better-than-rational outcomes can survive evolution in a sequential social dilemma. Using an evolutionary model of the Incremental Centipede Game, it shows that a level-k reasoning process whose errors systematically point toward later, higher-payoff nodes (positively biased reasoning) is consistently favoured by selection, while fully rational backward induction goes extinct. The bias co-evolves with bounded rationality: populations settle on shallow reasoning depths without any explicit cost for thinking. The paper also finds a stable coexistence between positively biased reasoners and myopic payoff-maximizers, and shows that longer games widen the conditions under which positive bias dominates. If correct, this turns a classic puzzle—why people deviate from game-theoretic rationality—into an adaptive outcome rather than a failure.","feed_headline":"Positive bias beats rational play in evolutionary centipede model","feed_subtitle":"Simulations of the centipede game show selection favors reasoning aimed at later, higher payoffs, while rational play dies out.","key_machinery":"The reasoning kernel matrix M(ε), a block matrix of noisy best-response probabilities for the two player roles. Applying this matrix k times to a myopic starting vector produces the full level-k strategy; its shape encodes the cognitive bias. The positively biased kernel B+ permits mistakes only in the direction of later nodes, the negatively biased kernel B− only toward earlier nodes, and the unbiased kernel spreads mistakes uniformly. This matrix is the mechanism linking a bias to an action distribution, and thus to payoffs and evolutionary success.","core_discovery":"In a six-step Incremental Centipede Game with exponentially growing stakes, each agent is defined by a starting behaviour and a reasoning kernel applied iteratively, giving σ(k) = σ(0)M(ε)^k. Three kernels are compared: unbiased reasoning, positively biased reasoning (deviations from the noisy best response go only toward later nodes, where payoffs are higher but uncertain), and negatively biased reasoning (deviations go only toward earlier nodes). Under selection in finite populations, the positively biased kernel becomes the most frequent reasoning type and the subgame-perfect equilibrium strategy essentially never invades. With the noise and selection strength calibrated to experimental d","pith_inferences":["The load-bearing part of the model is the precise shape of the biased kernels; alternative implementations—such as biasing beliefs about the opponent's reasoning depth instead of biasing action deviations, or making deviation probabilities decay with distance—could alter or overturn the dominance of positive bias.","The model implies a behavioural prediction: in centipede games with exponentially growing stakes, average stopping nodes should shift later as game length increases, beyond what unbiased level-k models predict.","The framework suggests that optimism in cognition may be a domain-specific response to asymmetric payoff growth rather than a general trait; games with shrinking future gains might instead favour negatively biased reasoning.","The predicted coexistence of myopic maximizers and optimistic one-step reasoners could be tested experimentally by classifying subjects' reasoning depth and checking whether the mixture matches the model's interior equilibrium."],"forward_implications":["Fully rational backward induction is not an evolutionary attractor once reasoning is noisy; the model predicts the subgame-perfect equilibrium strategy is driven to near extinction under strong selection.","Bounded rationality emerges without explicit cognitive costs: error propagation through iterative reasoning naturally limits the favoured reasoning depth.","Positively biased reasoning and no-reasoning can coexist stably, so populations can remain heterogeneous in sophistication even under strong selection.","Longer centipede games enlarge the cognitive-noise window in which positively biased reasoning is the majority strategy, linking the bias to the size of future gains.","The fitted model reproduces experimental stopping-node distributions, offering an alternative to explanations based on altruism or other-regarding preferences in the centipede game."],"supporting_citations":[{"why":"Introduces the centipede game, the sequential interaction setting in which the whole analysis is framed.","marker":"[25]"},{"why":"Provides the level-k reasoning framework that defines the recursive reasoning strategies of the model.","marker":"[21]"},{"why":"Supplies the noisy introspection model of stochastic deviations from best response that the reasoning kernels generalize.","marker":"[42]"},{"why":"Supplies the prior noisy level-k evolutionary model with shared reasoning noise that this paper extends by introducing distinct biased kernels.","marker":"[15]"},{"why":"Provides the pure-strategy evolutionary baseline for the centipede game whose conclusion—SPE prevails under strong selection—this model overturns.","marker":"[30]"},{"why":"Provides the American experimental centipede-game data used for the low-noise calibration of β and ε.","marker":"[2]"},{"why":"Provides the Japanese experimental centipede-game data used for the high-noise calibration of β and ε.","marker":"[22]"},{"why":"Supplies the replicator equation used to identify the stable coexistence equilibrium between positively biased reasoners and myopic maximizers.","marker":"[45]"},{"why":"Underpins the small-mutation limit approximation used to compute stationary distributions and invasion diagrams.","marker":"[65]"}],"fun_headline_variants":["Wishful thinking wins in centipede-game evolution","Rational strategy goes extinct in evolved centipede game","Evolution favors positive bias over cold logic in games","Biased reasoning outlives perfect play in simulation","Optimism evolves to survive social dilemmas"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The exact way the bias is built into the reasoning kernel—errors only toward later nodes for positive bias and only toward earlier nodes for negative bias, spread uniformly—carries the evolutionary outcome; a different implementation of the same bias could flip the result.","fun_headline_variants_meta":{"raw":{"variants":["Wishful thinking wins in centipede-game evolution","Rational strategy goes extinct in evolved centipede game","Evolution favors positive bias over cold logic in games","Biased reasoning outlives perfect play in simulation","Optimism evolves to survive social dilemmas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":2888,"prompt_tokens":742,"completion_tokens":2146,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":2073}},"tokens_in":486,"tokens_out":2146,"duration_ms":16380,"temperature":1.0,"reasoning_tokens":2073,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:48:12.936525+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same evolutionary dynamics with the positively biased kernel replaced by a kernel that biases beliefs about the opponent's reasoning level instead of biasing the action deviation; if positive bias no longer dominates, the result rests on the specific kernel shape. Behaviourally, the model predicts the noise window for majority positive bias widens with game length, so comparing stopping-node distributions across Incremental Centipede Games of lengths 4, 6, and 8 with identical per-step stakes would test the mechanism.","supporting_citations":[{"cited_title":"1981 Games of perfect information, predator y pricing and the chain-store paradox","cited_arxiv_id":null,"evidence_quote":"Introduces the centipede game, the sequential interaction setting in which the whole analysis is framed."},{"cited_title":"1993 Evolution of Smartn Players","cited_arxiv_id":null,"evidence_quote":"Provides the level-k reasoning framework that defines the recursive reasoning strategies of the model."},{"cited_title":"2004 A model of noisy introspection","cited_arxiv_id":null,"evidence_quote":"Supplies the noisy introspection model of stochastic deviations from best response that the reasoning kernels generalize."},{"cited_title":"2024 Evolut ion of a theory of mind","cited_arxiv_id":null,"evidence_quote":"Supplies the prior noisy level-k evolutionary model with shared reasoning noise that this paper extends by introducing distinct biased kernels."},{"cited_title":"2012 Evolutionary dynamics in ﬁnite popu- lations can explain the full range of cooperative behaviors observe d in the centipede game","cited_arxiv_id":null,"evidence_quote":"Provides the pure-strategy evolutionary baseline for the centipede game whose conclusion—SPE prevails under strong selection—this model overturns."},{"cited_title":"1992 An Experimental Study of the Cen - tipede Game","cited_arxiv_id":null,"evidence_quote":"Provides the American experimental centipede-game data used for the low-noise calibration of β and ε."},{"cited_title":"2012 Level-k analysis of experimental c en- tipede games","cited_arxiv_id":null,"evidence_quote":"Provides the Japanese experimental centipede-game data used for the high-noise calibration of β and ε."},{"cited_title":"1983 Replicator dynamics","cited_arxiv_id":null,"evidence_quote":"Supplies the replicator equation used to identify the stable coexistence equilibrium between positively biased reasoners and myopic maximizers."},{"cited_title":"2019 Computation and S imu- lation of Evolutionary Game Dynamics in Finite Populations","cited_arxiv_id":null,"evidence_quote":"Underpins the small-mutation limit approximation used to compute stationary distributions and invasion diagrams."}],"review_version":1}