{"id":"4a49d085-83f5-44e8-bc4b-cdf7c6dc38cb","arxiv_id":"2508.17151","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In a 14-parameter integrative experiment, punishment raised contributions almost universally but changed group welfare from about +43% to -44% depending on context, and an elastic-net model predicted new conditions better than human forecasters.","lead":"Researchers ran 360 public-goods-game conditions with 7,100 people, comparing games with and without punishment to see when punishment helps or hurts group welfare. Punishment consistently raised contributions, but its effect on earnings ranged from large gains to large losses depending on context, and a statistical model predicted those outcomes better than human experts and laypeople.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Platform confound in learning data may inflate attribution of predictive success to design features; validation R² and feature importance need robustness check.","rationale":"The reader's weakest_assumption identifies hidden moderators (platform, batch, chat content, recruitment pool) covarying with design parameters. The manuscript explicitly reveals a platform switch mid-way through the learning wave and a deterministic Sobol ordering, making this a real confound rather than a hypothetical one. This is the most load-bearing concern because it directly affects whether the model's predictive success and feature-importance rankings can be attributed to the 13 design parameters, which is the core of the central claim. The paper's strengths—the two-wave design, pre-registration, transparent SI, and a reproducible OSF repository—mean the issue is testable without new data collection. The failure of the expert-comparison CI to exclude zero is a secondary issue that can be fixed by softening the abstract; the hidden-moderator concern, if it lands, would require re-evaluating feature importance and the 'systematic patterns' interpretation. Therefore the verdict remains CONDITIONAL, with the platform-confound robustness check as the key uncertainty.","tokens_in":54566,"tokens_out":8390,"duration_ms":100861,"concrete_test":"Refit the complete E-Net pipeline (exactly as in SI S9: standardization, two-way interactions, Bayesian hyperparameter search with 50 iterations, 10-fold CV on the learning set) using only learning-wave experiments run on Prolific, i.e., excluding the 7 retained MTurk experiments. Then compute the out-of-sample R², RMSE, and the permutation importance increase for communication on the same 20 validation experiments. If the Prolific-only E-Net retains R² > 0.4 and communication remains the largest PFI with a >50% error increase, the platform confound is unlikely to drive the main conclusions. Additionally, report a logistic regression predicting MTurk vs Prolific from the 13 design parameters in the learning set to confirm whether platform is actually confounded with design features; if the model has low discriminative accuracy (e.g., AUC ≤ 0.7), the confounding channel is negligible.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that punishment-effectiveness heterogeneity is systematic and predictable from 13 PGG design parameters plus control efficiency. A load-bearing assumption is that these inputs exhaust the systematic determinants, so that the model's out-of-sample success and the feature-importance rankings (e.g., communication being three times more important than the next feature) can be attributed to the design parameters themselves. This assumption is threatened by the paper's own description in SI S5.2: the learning wave mixed 604 MTurk participants (from 7 early experiments retained after the platform switch) with 3,014 Prolific participants, while the validation wave was exclusively Prolific. Because the learning conditions were generated by a deterministic Sobol sequence and data collection proceeded in order, those 7 MTurk experiments are the earliest points in the sequence; the Sobol coordinates are therefore correlated with the platform/time period. If MTurk participants differed systematically (e.g., in attention, comprehension, or punishment norms), the model could learn associations between design parameters (or their interactions) and outcomes that are actually platform effects. Such spurious associations would not transfer to the Prolific validation set, so the reported R² = 0.53 might understate true design-feature predictability, but the contribution of individual design features—especially communication, which may coincide with the later Prolific phase—could be overestimated. The manuscript's statement that 'this heterogeneity can only be attributed to differences in game parameters' (Results) is thus an overreach, and the authors even concede in the Discussion that their prediction exercise is 'likely an optimistic assessment' because population variation is unexplored. The platform mix does not invalidate the prediction exercise, but it is a concrete uncontrolled moderator that undermines the strong attribution of patterns to design parameter","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale integrative experiment in which 14 public-good-game (PGG) design parameters are systematically varied, using 320 Sobol-sampled conditions in a learning wave and 40 pre-registered validation conditions with 8–12 replicates each (360 conditions total, 147,618 decisions). The central finding is that punishment reliably increases contributions but has highly heterogeneous effects on normalized efficiency, ranging from a 43% improvement to a 44% reduction across settings. The authors train several machine-learning models on the learning wave to predict treatment-group efficiency from the 13 design parameters plus control-group efficiency, and the elastic-net model achieves an out-of-sample R² of 0.53 on the 20 validation experiments, outperforming 553 human forecasters. Feature-importance analyses identify communication as the dominant predictor (about three times the next feature), with important interactions involving game length, contribution framing, contribution type, and peer-outcome visibility. The paper argues that heterogeneity in punishment effectiveness is systematic and predictable, rather than noise, and that the debate should shift from whether punishment works to the conditions under which it helps or harms welfare.","tokens_in":54853,"tokens_out":4585,"duration_ms":61395,"significance":"If the findings are robust, this is a substantial contribution: it is among the largest single experimental studies of punishment, combines a space-filling design with a pre-registered out-of-sample validation phase, provides open data and code, and demonstrates that a simple regularized model can outperform human experts at predicting the welfare consequences of punishment in new settings. The design's two-wave structure and the use of permutation tests and I² statistics are strengths. The paper also makes the useful methodological point that single-factor studies cannot quantify relative feature importance or detect higher-order interactions. However, the load-bearing inference that heterogeneity is attributable to design parameters, and the specific claim that communication is the dominant determinant, rest on an assumption of uniform protocols across learning and validation waves that is contradicted by the paper's own platform-switch description (SI S5.2). Because this assumption is central to the interpretation of both the out-of-sample R² and the feature-importance rankings, it must be addressed with a targeted robustness analysis before the paper's primary claims can be ful","major_comments":[{"comment":"The paper claims in Results that \"because all experiments were conducted under identical protocols, recruitment procedures, and interfaces, this heterogeneity can only be attributed to differences in game parameters.\" This is contradicted by SI S5.2: the learning wave mixed 604 MTurk participants (7 early experiments) with 3,014 Prolific participants, while the validation wave was entirely Prolific. Because the learning conditions were generated sequentially by a Sobol sequence, the 7 MTurk experiments are the earliest Sobol points, so platform/phase is correlated with design coordinates. The prediction model (Methods, Model Construction) is trained on D_learning, which includes these 7 MTurk pairs, and then validated on Prolific-only data. If MTurk participants differ in attention, comprehension, or behavioral norms, the model could learn spurious associations between design parameters","section":"SI S5.2 and Results (\"Heterogeneous effects...\")"},{"comment":"The headline out-of-sample R² of 0.53 is computed on 20 validation experiments. The paper provides bootstrap 95% CIs for RMSE (SI Table S8; E-Net RMSE CI [3.35, 5.69], ratio to baseline [0.53, 0.90]), but not for R². Since R² is a nonlinear transformation of the same residuals, the RMSE CI does not directly convey the uncertainty of R². Given the small number of validation conditions, the R² estimate could be imprecise; a bootstrap CI for R² (or the ratio of mean squared errors) should be reported. This is load-bearing because the abstract and discussion repeatedly emphasize R²=0.53 as the quantitative evidence that the heterogeneity is systematic and predictable.","section":"Results (\"Predicting punishment effectiveness...\") and SI S9.3"}],"minor_comments":[{"comment":"The abstract says \"sampling 360 experimental conditions\" but the Introduction says 320 learning conditions and 40 validation conditions; after filtering, the learning wave is described as 335 games / 150 matched pairs. Please make the count of conditions vs. games vs. matched pairs consistent and explicit.","section":"Abstract / Introduction"},{"comment":"The text says \"13 other PGG design parameters are used as inputs\" because reward impact is dropped; this is fine but should be stated in the main text (Table 1 lists 14 parameters) so readers know why reward impact is excluded. Also specify which of the 13 are binary vs. continuous.","section":"Methods, Model Construction"},{"comment":"The paper honestly reports that the model hyperparameters were not pre-registered in advance of collecting the training data; the \"stricter pre-registration strategy\" note is appropriate, but the main text's characterization of the two-preregistration procedure should be softened to avoid implying the entire modeling pipeline was pre-registered.","section":"SI S1"},{"comment":"Please report how many of the 150 matched learning pairs involve MTurk participants, and their design-parameter coordinates, so readers can assess the degree of correlation between platform and the Sobol sequence. This is directly relevant to the first major comment.","section":"SI S5.2"},{"comment":"The caveat that the prediction exercise \"likely represents an optimistic assessment of performance\" is a useful limitation statement, but it is positioned as a general point about population variation. It should be moved closer to the results or methods and explicitly connected to the platform switch in the learning wave.","section":"Discussion (penultimate paragraph)"}],"recommendation":"major_revision","confidential_remarks":"The platform confound in the learning data is the main substantive issue. It is not a fatal flaw, but it must be addressed with a Prolific-only robustness analysis; if the feature importance rankings survive that check, the paper's central claims are supported. The authors should also be encouraged to report a bootstrap CI for R², as the current RMSE CI does not fully capture the uncertainty of the headline quantity. The dataset and the two-wave design are genuinely valuable, and the comparison to human forecasters is compelling. Overall, this is a well-executed and ambitious study that merits publication after the requested robustness work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious paper and the first to simultaneously vary 14 PGG parameters at scale and show out-of-sample prediction of punishment effectiveness. The two-wave design is a real strength: Wave 1 Sobol-sampled learning data, Wave 2 pre-registered validation with 8–12 trials per condition, open code and data. The E-Net R² of 0.53 on held-out conditions is real evidence that the heterogeneity is systematic, not noise, and the feature-importance finding that communication dominates is interesting and consistent with prior single-factor work. No circularity: models are fit on Wave 1 and tested on Wave 2; using control efficiency as an input is legitimate, not leakage of the target.\n\nMain soft spots are three. First, the learning wave mixed 604 MTurk participants with 3,014 Prolific, and the MTurk sessions are the earliest in the Sobol sequence, so design parameters and platform/time are correlated. That makes the attribution of predictive success to specific design features—especially the \"communication is three times more important\" claim—less clean. The validation wave is Prolific-only and still predicts well, so the central claim survives, but the authors should re-estimate without the MTurk games or include platform as a feature to show feature importances are stable. Second, the abstract says the models \"outperformed\" domain experts; the bootstrapped CI for the expert difference in RMSE is [-0.10, 3.71], which includes zero, so the wording overstates the evidence. Third, the single reward-ring game (efficiency 27) is excluded as an outlier; that is defensible but post-hoc, and it should be shown as a sensitivity rather than quietly filtered.\n\nNone of these break the main argument. The paper delivers a useful quantitative map of when punishment helps or hurts welfare, and it demonstrates the integrative-experiment-plus-ML recipe cleanly. The authors even concede in the Discussion that their prediction performance is likely optimistic because population variation is unexplored, which is the right honest framing.\n\nThis deserves a serious referee. I'd send it to review and ask for the platform-robustness check and a softer abstract before final acceptance. It's the kind of paper I'd assign in a reading group on experimental methods and prediction.","headline":"A solid, high-scope punishment experiment with genuine out-of-sample prediction; needs two caution flags on platform confound and the expert-comparison claim before its headline statements are taken at face value.","tokens_in":55430,"tokens_out":2651,"would_cite":true,"duration_ms":32986,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Punishment consistently raises contributions to public goods, but across 360 game configurations its effect on welfare swings from a 43% gain to a 44% loss, and a model trained on those designs predicts outcomes in new experiments better th","keywords":["public goods games","punishment","cooperation","social dilemmas","integrative experiments","machine learning prediction","efficiency","communication"],"falsifier":"Run the same trained models on a fresh, pre-registered wave of randomly sampled conditions drawn from the same 14-dimensional space, with participants from a different recruitment source and sessions run at different times; if out-of-sample R² falls to zero or below the human benchmark, the claim that the design parameters are the systematic drivers fails.","tokens_in":54453,"feed_emoji":"💰","tokens_out":5887,"duration_ms":67753,"temperature":0.7,"pith_summary":"Punishment reliably raises how much people contribute to a public good, but whether that helps or hurts the group's overall earnings is highly context-dependent: across 360 sampled game configurations, punishment improved welfare by as much as 43% and reduced it by as much as 44%. The paper argues this spread is systematic, not noise, because a statistical model trained on design parameters and control-condition efficiency predicted punishment outcomes in 20 brand-new experiments with out-of-sample R² of 0.53, while averaged predictions from domain experts and laypeople were close to no better than guessing. Communication was the single most important design feature, followed by contribution framing (opt-out vs. opt-in), contribution type, game length, visibility of others' earnings, and the presence of rewards. Most of these features combine through interactions rather than as independent add-ons—for example, longer games only help punishment when the group can communicate. The upshot is a shift from asking whether punishment 'works' to asking under which specific conditions it works.","feed_headline":"Punishment can lift welfare 43% or slash it 44%","feed_subtitle":"A model trained on 360 experimental settings predicts when punishment helps, and communication is the top lever.","key_machinery":"The central object is a 14-dimensional design space of public goods games, with each condition run as a matched pair (punishment enabled vs. disabled). The carrying mechanism is an elastic-net regression with pairwise feature interactions that takes the 13 design parameters plus the no-punishment control efficiency as inputs and predicts the punishment treatment efficiency; it is interpreted with permutation importance and SHAP values, and validated on pre-registered out-of-sample conditions. The paper also introduces 'normalized efficiency' as a scale-free outcome so game configurations with different multipliers and group sizes can be compared.","core_discovery":"The central claim is that punishment's welfare effect is governed by specific, discoverable combinations of the game's design parameters. Averaged across all conditions, punishment hurt efficiency (from 0.71 to 0.63 in the learning wave), but this average masks large and statistically significant variation, with some configurations seeing welfare increase by roughly 43% and others fall by roughly 44%. The authors show that a model can predict these outcomes out-of-sample, which implies the heterogeneity is predictable, and that the dominant predictor is whether groups can communicate, with opt-out contribution framing, all-or-nothing vs. variable contributions, game length, and visibility of","pith_inferences":["If the model's success generalizes beyond online Western participant pools, the same design space could be used to map where punishment-based institutions will backfire before they are deployed in field settings.","Because communication dominates, an untested extension is whether the content of chat matters—whether norm-setting and coordination, rather than mere presence of a chat window, is what drives the interaction with game length.","The finding that punishment-technology parameters barely predict outcomes suggests that efforts to tune the severity or cost of sanctions may matter far less than the social and informational context in which they are embedded—an inference the paper only partially endorses.","A direct test of the communication mechanism: add a condition where communication is enabled but restricted to non-task chat; if communication's importance persists, the mechanism is not coordination but something like group identity."],"forward_implications":["Policy and design decisions about introducing peer punishment should be made condition-by-condition, not on the average effect, because the same mechanism that raises contributions can either roughly double or roughly halve welfare.","Communication is not just another factor; it is the dominant lever, three times more important than any other parameter, and it changes whether longer interactions help or hurt.","Rarely studied features like the opt-out default matter more than the punishment technology's cost and impact, redirecting research attention from punishment mechanics to defaults, framing, and information structure.","The two-wave design plus out-of-sample prediction provides a template for turning heterogeneous, seemingly contradictory experimental results into quantitative predictions in other contested domains.","Because most important features interact, single-factor studies and meta-analyses that pool across unspecified procedures will keep producing conflicting answers unless the design space is sampled systematically."],"supporting_citations":[{"why":"Foundational demonstration that costly punishment increases contributions, the baseline effect this paper extends and qualifies.","marker":"Fehr & Gächter, 1999, 2002"},{"why":"Counter-result showing punishment can reduce welfare; the heterogeneity this paper systematically explains.","marker":"Dreber et al., 2008"},{"why":"Evidence of antisocial punishment across societies, motivating the search for context-dependent effectiveness.","marker":"Herrmann et al., 2008"},{"why":"Meta-analysis of reward and punishment providing the comparator for scale and the reward-versus-punishment debate.","marker":"Balliet et al., 2011"},{"why":"Shows communication plus punishment in voluntary contribution games, the empirical basis for the top-ranked feature.","marker":"Bochet et al., 2006"},{"why":"Demonstrates long-run benefits of punishment, the result that the game-length interaction extends and conditions.","marker":"Gächter et al., 2008"},{"why":"Supplies the integrative-experiment design rationale of sampling many conditions to estimate interactions.","marker":"Almaatouq et al., 2022"},{"why":"Provides the elastic-net regularized regression method used to build the predictive model.","marker":"Zou & Hastie, 2005"}],"fun_headline_variants":["Punishment boosts welfare 43% or hurts 44% depending on context","Communication is top factor in when punishment helps or hurts","Model predicts punishment's welfare swings from -44% to +43%","Punishment's impact varies: +43% to -44% welfare change","When punishment works: context matters more than presence"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The prediction result assumes that the 13 game-design parameters plus the no-punishment baseline exhaust the systematic causes of punishment effectiveness, so if hidden factors—like which recruitment panel participants came from, when the session ran, or what was said in chat—vary with the design, the model's out-of-sample accuracy could be partly an artifact of those correlates.","fun_headline_variants_meta":{"raw":{"variants":["Punishment boosts welfare 43% or hurts 44% depending on context","Communication is top factor in when punishment helps or hurts","Model predicts punishment's welfare swings from -44% to +43%","Punishment's impact varies: +43% to -44% welfare change","When punishment works: context matters more than presence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1333,"prompt_tokens":818,"completion_tokens":515,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":424}},"tokens_in":562,"tokens_out":515,"duration_ms":5996,"temperature":1.0,"reasoning_tokens":424,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:59:29.065643+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same trained models on a fresh, pre-registered wave of randomly sampled conditions drawn from the same 14-dimensional space, with participants from a different recruitment source and sessions run at different times; if out-of-sample R² falls to zero or below the human benchmark, the claim that the design parameters are the systematic drivers fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the elastic-net regularized regression method used to build the predictive model."}],"review_version":1}