{"id":"9e130032-17a0-4f54-a840-91a7179571b3","arxiv_id":"2608.01183","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In a four-strategy Prisoner's Dilemma with peer punishers and rewarders, institutional rewards should target the enforcers, institutional punishment should target defectors only, and punishing enforcers destroys cooperation and welfare.","lead":"This paper models how a central institution should reward or punish different strategies in a society that already has its own peer punishment and peer reward. It finds that subsidizing ordinary cooperators does little, while rewarding the people who enforce norms works best, and punishing those enforcers backfires.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Welfare rankings rest on unspecified per-capita vs per-interaction cost accounting: Eq. (4) adds theta_i to every interaction while Section 3.6 charges |theta_i| once per capita, so the policy rankings in Figs. 9 and 14 are not well-defined as written.","rationale":"The reader and I identify the same weakest point: institutional cost accounting. I think it is even more severe than 'real institutions may face convex costs': the paper does not define a computable welfare function, and the implementation of theta_i in Eq. (4) as a per-interaction payoff shift is incompatible with the per-capita cost stated in Section 3.6. Because the paper's contribution is a policy ranking by social welfare, this is load-bearing. The cooperation-frequency claims in Figures 5, 6, 10, and 11 may survive, but the design principles in the abstract and Discussion, such as 'institutions should prioritise penalising defectors' and 'use targeted rewards for social rewarders', are welfare claims and cannot be evaluated as written. I would keep the reader's CONDITIONAL verdict, with the condition being a consistent welfare definition and re-derived rankings. I also note the paper itself flags a second limitation in Figure 11, where punishment has no effect at low delta, which undercuts the phrase 'consistently effective'; however, the welfare accounting issue is more fundamental because it undermines every welfare comparison in the paper.","tokens_in":17902,"tokens_out":7200,"duration_ms":71012,"concrete_test":"Ask the authors to write the exact welfare formula used to compute Figures 9 and 14, then recompute those figures under the accounting implied by Eq. (4): charge 4*|theta_i| per targeted individual per round (degree-matched cost) while keeping the 4*theta_i payoff benefit; and under the alternative transfer convention where theta_i is removed from player payoffs and only the cost is subtracted. If the relative ranking of {SP,SR} over {C} under reward, or of {D} over {SP} under punishment, changes between these conventions, the headline design principle is an artifact of the accounting choice and must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central design principles are welfare claims, but the welfare measure is never stated and the cost accounting implied by the model is internally inconsistent. In Section 3.3, the institution 'pays |theta_i| per targeted individual to shift strategy i's payoff by theta_i', and Eq. (4) implements this as a row-wise addition of theta_i to every entry of the per-interaction payoff matrix P. On the L=100 lattice each player accumulates payoffs from four neighbours, so a targeted player receives 4*theta_i per round, while Section 3.6 describes 'per capita cost value theta=1.0' and an 'efficiency coefficient a=1' that never appears in any formula. If welfare is aggregate payoff net of cost, a reward policy can create 3*theta_i of net value per target per round purely by the interaction-count mismatch; if instead the intent is a once-per-capita transfer, Eq. (4) overstates the payoff shift. Under the natural alternative in which institutional transfers are excluded from players' payoffs and only the cost is subtracted, every reward and punishment policy becomes a net cost, potentially reversing the ranking that rewards SP/SR and punishes D. The paper's own Figure 11 already shows all punishment policies are indistinguishable at delta=0.4 and delta=1, so the 'consistently effective' claim is also narrower than stated, but the welfare accounting is the more fundamental problem because it affects every welfare comparison in Figures 9 and 14.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a four-strategy Prisoner's Dilemma in which unconditional cooperators (C), defectors (D), social punishers (SP), and social rewarders (SR) coexist, and an external institution can add a fixed payoff shift theta_i to any subset of strategies. The authors derive replicator dynamics and equilibrium conditions for infinite well-mixed populations, and run agent-based simulations on a 100x100 lattice with a Fermi update rule. They report that peer punishment promotes cooperation most strongly while peer reward is better for welfare; that institutional rewards should target SP and SR rather than C; and that institutional punishment should target D only, since punishing SP or SR destroys cooperation and welfare. The paper's central design principles are welfare-based.","tokens_in":18264,"tokens_out":9843,"duration_ms":77634,"significance":"If the welfare rankings were rigorously established, the paper would offer a useful design principle for institutional incentive schemes in social dilemmas, and the four-strategy model is a natural extension of earlier three-strategy analyses. Strengths include analytic equilibrium characterization for the well-mixed case, parameter sweeps over (epsilon,delta) for all target subsets, and specific falsifiable predictions about which strategies institutions should reward or punish. However, the central welfare measure is not fully specified and the analytical section contains internal contradictions, so the significance is conditional on fixing these issues.","major_comments":[{"comment":"The welfare measure used for the central policy rankings is never written down, and the implied cost accounting is internally inconsistent. Eq. (4) implements theta_i as a payoff shift added to every entry of row i, so on the L=100 lattice each targeted player receives 4*theta_i per round from its four neighbors. Section 3.6, however, describes a \"per capita cost value theta=1.0\" and an efficiency coefficient a=1 that appears in no formula. If welfare is aggregate payoff net of a once-per-capita cost |theta_i|, then any reward policy creates 3*theta_i of apparent net value per target per round purely from the interaction-count mismatch; if instead the transfer is meant to be netted out of payoffs, Eq. (4) overstates the payoff shift and every policy becomes a net cost. Either way, the policy rankings in Figs. 9 and 14 are not well-defined as written, and the authors must state the exact welfare function, specify whether theta_i is per interaction or per capita, and re-run or re-derive the comparisons.","section":"Sections 3.3 and 3.6, Eq. (4), Figs. 9 and 14"},{"comment":"The stability statements in Section 4.1.3 and the captions of Figures 3 and 4 are mutually contradictory. Figure 3's caption says \"no stable equilibrium exists\" while also saying streamlines converge toward C, and the text then calls vertex C unstable. If trajectories converge to C, then C is at least attracting, so calling it unstable is not coherent. Likewise, Figure 4's caption says D is stable except for SP in the (0,0,+,+) and (0,0,+,-) cases, while the text says \"The equilibrium at vertex D is consistently stable through all 4 policy regimes.\" These contradictions concern the core analytical claim about which equilibria are stable and should be resolved with a correct stability classification for each vertex and for the reported interior and edge equilibria.","section":"Section 4.1.3 and Figures 3 and 4"},{"comment":"The claimed continuum of equilibria on the C-D edge in Figure 3 is inconsistent with the paper's own edge analysis. Section 4.1.2 gives the C-D edge interior point as y=(R-T-gamma)/B. With the Figure 3 parameters theta_C=1, theta_D=0 (so gamma=-1) and the stated payoff values R=3, T=5, S=0, P=1, one obtains B=-1 and y=1, which is not an interior point of the edge. The edge flow is therefore monotone rather than containing a continuum of equilibria, so the caption's claim that a continuum exists \"consistent with the theoretical stability analysis\" needs either correction or a detailed derivation.","section":"Section 4.1.2 and Figure 3 caption"},{"comment":"The headline claim that penalising defectors is \"the only consistently effective policy\" is stronger than the presented evidence. Figure 11 shows that at delta=0.4 and delta=1 the eight punishment policies are indistinguishable and all decay to zero, with separation only at delta=3. The text itself acknowledges that \"the policy ranking of Figure 10 emerges only once peer incentives are strong enough to matter.\" This parameter dependence should be stated prominently, and the abstract and discussion should not present the ranking as uniform over the (epsilon,delta) regimes examined.","section":"Section 4.2.2 and Figure 11"}],"minor_comments":[{"comment":"The text refers to \"Figure 1\" when reporting stationary cooperation levels, but the relevant figure appears to be Figure 5; several other cross-references need checking.","section":"Section 4.2.1"},{"comment":"The efficiency coefficient a=1 is introduced but never appears in any formula; either define how a enters the welfare calculation or remove it.","section":"Section 3.6"},{"comment":"The sentence stating that the choice of target set is inert in the delta-close-to-epsilon regime appears twice in consecutive paragraphs; delete the duplicate.","section":"Section 4.2.2, around Figure 12"},{"comment":"The paper should state explicitly whether \"cooperation level\" counts SP and SR as cooperative strategies, since the figures plot combined C+SP+SR frequency while the discussion sometimes contrasts SP/SR with plain C.","section":"Figures 6 and 11"},{"comment":"There are minor language errors, for example \"may not be captures\" in Section 1 and \"subsiding SP\" in Section 4.1.3 (likely \"subsidising\"); a careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The welfare-accounting problem is the most serious issue because welfare is the stated basis for the paper's design principles; the manuscript cannot be accepted until the welfare function is explicit and internally consistent. If the corrected welfare definition reverses the policy rankings, the headline conclusions will change. The contradictory stability statements in Section 4.1.3 and Figures 3-4 also need to be resolved before the analytical part can be trusted. The paper would also benefit from a data/code availability statement given the simulation-based nature of the results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the read worth having here is the four-strategy extension: C, D, SP, and SR with institutional incentives that can target any subset. The exhaustive comparison of seven reward and seven punishment target sets on an (epsilon,delta) grid is systematic, and the core qualitative message—reward the enforcers, punish defectors, never punish SP or SR—is plausible and broadly consistent with prior work. The analytical section also does real work: the interior equilibrium is shown to be a one-parameter family under a solvability condition, and the edge/face analysis is careful. The heavy self-citation is not by itself a problem; the cited prior work is relevant.\n\nWhere the paper is soft, and these are not minor: first, the welfare measure is never written down. Section 3.3 says the institution pays |theta_i| per targeted individual to shift strategy i's payoff by theta_i, but Eq. (4) adds theta_i to every row entry, i.e. every pairwise interaction. On the L=100 lattice a targeted player gets 4*theta_i per round while being charged theta_i once. That mismatch can manufacture net welfare gains out of thin air, and it undermines every ranking in Figures 9 and 14. The efficiency coefficient a=1 is mentioned once and never defined. Second, the analytical text contradicts itself: Figure 3 says no stable equilibrium exists and streamlines converge to C, then calls C unstable; Figure 4 says D is stable except for the SP cases, then says D is stable through all four regimes. Third, the simulations report one average over the last 1000 steps of what appears to be a single run—no error bars, no seeds, no code. For a paper whose central claims are quantitative welfare comparisons, that is thin. Fourth, the phrase 'only consistently effective policy' is too strong given Figure 11, where all punishment policies are indistinguishable at delta=0.4 and delta=1.\n\nThe stress-test note is right: the per-capita versus per-interaction cost accounting is the load-bearing problem. If the institution's transfer enters the players' payoffs per interaction but is costed per capita, the welfare comparisons are not well-defined as written. That said, the underlying model is not absurd, and the qualitative insight about sanctioning enforcers is worth testing.\n\nWho should read this? Groups working on institutional incentive design in evolutionary games, and anyone building on welfare-oriented incentives. It deserves a serious referee, but with the expectation of major revision: the welfare accounting must be fixed and stated, the contradictions resolved, and the code released. I would not cite the design principles until that happens.","headline":"A useful four-strategy extension with plausible design principles, but the welfare rankings are not well-defined as written because of per-capita vs per-interaction cost accounting, plus internal contradictions in the stability text.","tokens_in":18813,"tokens_out":3218,"would_cite":false,"duration_ms":29202,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A22","91A80"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under institutional punishment, only punishing defectors works; under institutional reward, institutions should target the peer enforcers, because sanctioning punishers or rewarders destroys cooperation and welfare.","keywords":["evolution of cooperation","social dilemma","peer punishment","peer reward","institutional incentives","social welfare","replicator dynamics","agent-based simulation"],"falsifier":"In a laboratory or agent-based replication, impose an institution that fines defectors and also fines peer punishers at the same fixed per-capita budget; the model predicts cooperation and welfare fall below the no-institution baseline whenever peer punishers are targeted. If a well-powered replication instead finds cooperation maintained or welfare higher, the claim that sanctioning enforcers is always harmful would be falsified.","tokens_in":17674,"feed_emoji":"⚖️","tokens_out":7027,"duration_ms":58844,"temperature":0.7,"pith_summary":"The paper asks which institutional interventions work when individuals can already reward and punish each other in a Prisoner's Dilemma with four strategies: unconditional cooperators, defectors, social punishers, and social rewarders. An external institution may reward or punish any subset of these four types, and the paper measures success by both cooperation level and social welfare, defined as aggregate payoff minus the institution's per-capita cost. Using replicator dynamics for infinite well-mixed populations and agent-based simulations on a square lattice, it finds a sharp asymmetry: institutional reward works best when aimed at the peer-incentive strategies (especially social rewarders), while subsidizing plain cooperators has little effect; institutional punishment works only when aimed at defectors, and punishing peer enforcers destroys cooperation and welfare. The paper concludes that institutions should never sanction social punishers or rewarders, should reward rewarders, and should spend punishment budget on defectors. Maximizing cooperation alone is a misleading objective because punishment-heavy regimes can yield low welfare.","feed_headline":"Punish only defectors, reward the enforcers","feed_subtitle":"Penalizing peer punishers breaks enforcement; rewarding enforcers builds welfare and cooperation.","key_machinery":"The argument is carried by a 4x4 payoff matrix in which social punishers reduce a defector's payoff by $\\delta_P$ at personal cost $\\epsilon_P$ and social rewarders increase a cooperator's payoff by $\\delta_R$ at cost $\\epsilon_R$, with the institution applying a fixed additive payoff shift $\\theta_i$ to every payoff of each targeted strategy $i$ (positive reward, negative punishment) at cost $|\\theta_i|$ per targeted individual. Welfare is aggregate population payoff minus that institutional cost, and the evolutionary dynamics are the replicator equations in both well-mixed and lattice populations, with the lattice using a Fermi imitation rule. A structural feature drives the analysis: with institutional shifts, the interior equilibrium exists only under a solvability condition $\\alpha/\\epsilon_P + \\beta/\\epsilon_R = 1$ relating the reward shifts to the peer costs, and when it exists it is a one-parameter line segment rather than an isolated point, meaning the institution chooses where on that line the population coexists. The mechanism behind the main result is that SP and SR are themselves cooperators, so punishing them both removes cooperative strategies and dismantles the peer-enforcement channel, whereas punishing D directly makes cooperation viable without needing peer enforcement.","core_discovery":"In a four-strategy Prisoner's Dilemma where social punishers and social rewarders coexist with unconditional cooperators and defectors, the effect of an institutional intervention depends almost entirely on which subset of strategies it targets. When the institution rewards, the only policies that substantially raise both cooperation and social welfare are those that include the peer-incentive strategies SP and/or SR, with {SP, SR} performing best; rewarding C alone produces cooperation levels indistinguishable from the no-institution baseline and adding C to a target set weakens the policy. When the institution punishes, the ranking reverses: punishing D alone is the only consistently effective policy, while punishing SP or SR dismantles the decentralized enforcement that keeps defection in check and drives cooperation below the baseline. Punishing enforcers is not merely ineffective but destructive. Throughout, defectors persist in essentially all parameter regions, and the institution mainly reshapes which cooperative and enforcer types coexist with defection and at what welfare level; peer punishment is the strongest promoter of cooperation, while peer reward is the better guardian of social welfare.","pith_inferences":["The interior equilibrium is a line segment, so an institution's budget can be viewed as selecting a point on a coexistence continuum rather than forcing a single winning strategy; this suggests fine-tuned policies could park a society near a welfare-optimal mix without eliminating defection.","The result that adding unconditional cooperators to a reward target weakens the policy implies that 'cooperator' is not a homogeneous category for policy design; institutions should distinguish plain cooperators from peer enforcers.","Because punishing SR is roughly twice as harmful to welfare as punishing SP, a testable design rule extends beyond the paper: among enforcer types, rewarders deserve stronger institutional protection than punishers since they generate positive-sum spillovers.","A natural experimental test follows from the model: a lab public-goods game with an external fine on vigilantism should reproduce the collapse in cooperation, and an explicit external subsidy of peer rewarders should raise welfare more than an equal subsidy of ordinary cooperators."],"forward_implications":["Under institutional punishment, the budget should be spent on D only: any policy that also targets SP or SR lowers both cooperation and welfare, and targeting enforcers without D approaches uniform defection.","Under institutional reward, the effective targets are SP and SR, not C: {SP, SR} yields the highest cooperation and welfare, and adding C to a reward set reduces the policy's effect.","Evaluating an institution by cooperation alone is misleading: the same intervention that maximizes cooperation can be welfare-negative, and peer reward outperforms peer punishment on welfare while trailing it on cooperation.","In structured populations the same policy can produce different coexistence outcomes than in well-mixed populations, so the design rules depend on the interaction network.","Defectors persist in essentially every scenario; institutions should think of their role as shaping which cooperative and enforcing strategies coexist with defection."],"supporting_citations":[{"why":"The three-strategy peer-incentive model this paper extends to four.","marker":"Han, 2016"},{"why":"Classic reward-and-punishment baseline for peer incentives.","marker":"Sigmund et al., 2001"},{"why":"Spatial public goods game with reward; precursor for network effects.","marker":"Szolnoki and Perc, 2010"},{"why":"Cost-effective external interference framework for institutional design.","marker":"Han and Tran-Thanh, 2018"},{"why":"Cost-efficiency analysis of institutional incentives that motivates the cost-accounting welfare measure.","marker":"Duong and Han, 2021"},{"why":"Local and global interference in structured populations; supplies the network-intervention approach.","marker":"Han et al., 2018"},{"why":"Defines cooperation versus social welfare as the evaluation lens used here.","marker":"Han et al., 2025"},{"why":"Provides lattice-population update rules (Fermi imitation) for the agent-based simulations.","marker":"Szabó and Fath, 2007"}],"fun_headline_variants":["Reward enforcers, punish only defectors","Punishing enforcers backfires; reward them instead","Institutional design: punish defectors, reward enforcers","Target enforcers for reward, defectors for punishment","Reward punishers and rewarders, punish only defectors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The policy rankings assume an institution can shift each targeted strategy's payoff by a fixed per-person amount $\\theta_i$, costs $|\\theta_i|$ per targeted individual, and that social welfare equals aggregate payoff minus that per-capita cost; if enforcement costs are convex, charged per interaction, or change how incentives combine with peer effects, the ranking of policies could change.","fun_headline_variants_meta":{"raw":{"variants":["Reward enforcers, punish only defectors","Punishing enforcers backfires; reward them instead","Institutional design: punish defectors, reward enforcers","Target enforcers for reward, defectors for punishment","Reward punishers and rewarders, punish only defectors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1472,"prompt_tokens":992,"completion_tokens":480,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":397}},"tokens_in":608,"tokens_out":480,"duration_ms":3964,"temperature":1.0,"reasoning_tokens":397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:11:02.705707+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a laboratory or agent-based replication, impose an institution that fines defectors and also fines peer punishers at the same fixed per-capita budget; the model predicts cooperation and welfare fall below the no-institution baseline whenever peer punishers are targeted. If a well-powered replication instead finds cooperation maintained or welfare higher, the claim that sanctioning enforcers is always harmful would be falsified.","supporting_citations":[],"review_version":2}