{"id":"93cb3534-3d5f-408b-ba6b-086ebd2419d1","arxiv_id":"2607.03113","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Under voluntary opt-in, matching outperforms price-equivalent rebate nearly fourfold because of advantageous selection into matching and disadvantageous selection into rebate.","lead":"A nationwide experiment with 2,400 Japanese adults finds that matching and rebate subsidies look equivalent under compulsory assignment once budgets and comprehension are fixed, but under voluntary opt-in the matching advantage nearly quadruples because rebate loses power. Self-selection, not price, can reverse which instrument works better in practice.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged external-validity limit; the experimental claim holds under the paper's own design.","rationale":"The paper's contribution is the isolation of the voluntary-take-up margin under otherwise refined conditions. Inside that design the numbers line up: compulsory near-equivalence after budget and comprehension adjustments, re-emergence under opt-in driven by rebate attenuation, and asymmetric LATEs. The reader's CONDITIONAL verdict already correctly weights the external-validity gap (procedural burdens, timing, awareness) and the secondary issues of non-pre-specification and missing code. No additional load-bearing internal flaw—budget mis-equalization, LATE identification failure under the Fowlie et al. (2021) decomposition, or contradiction with the estimated γ≈0—undermines the experimental claim. Therefore the verdict remains CONDITIONAL; the concrete check simply verifies that the headline pattern is not an artifact of the high-comprehension cut.","tokens_in":30450,"tokens_out":550,"duration_ms":6549,"concrete_test":"Re-estimate Table 3 Column 2 and the TOT/TOU decomposition of Figure 2 on the full (unrestricted) sample and on the low-comprehension subsample separately; if the opt-in matching-minus-rebate gap remains positive and statistically significant and the LATE asymmetry (TOT>ATE matching; TOT≤ATE rebate) is directionally preserved, the experimental claim is robust to the non-pre-specified sample cut.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central experimental claim is internally well-supported: after budget equalization (Section 3.4, Table C.1) and the high-comprehension cut (Table 2 cols 5–6), the compulsory gap is 26 JPY (p=0.351); under opt-in the ITT gap expands to 97.7 JPY (Table 3) via rebate ITT collapse and matching ITT stability, with the LATE ordering TOT>ATE for matching and TOT≤ATE for rebate (Figure 2). The reader's weakest assumption—that a symmetric low-cost default-off click isolates the real-world selection margin—is already the binding external-validity caveat (Section 6, Eckel-Grossman 2017 gap). No stronger internal inconsistency, identification failure, or arithmetic error overturns the reduced-form result inside the experiment. The non-pre-specified comprehension restriction and exploratory consumption-motive analysis (Section 5.3) are transparently disclosed and do not reverse the ranking.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper tests whether the refined near-equivalence of 1:1 matching and 50% rebate subsidies under compulsory assignment survives voluntary take-up. In a pre-registered, incentivized online experiment with a quota-balanced sample of 2,400 Japanese adults, the authors cross subsidy type with assignment rule (compulsory vs. opt-in), equalize budget constraints via dual upper limits, and restrict attention to high-comprehension participants. Under compulsory assignment the gap in total amount received by the charity shrinks to a statistically insignificant 26 JPY; under opt-in it re-emerges at 97.7 JPY (nearly four times larger) because rebate ITT falls by roughly half while matching ITT remains near its ATE. LATE decompositions following Fowlie et al. indicate advantageous selection into matching (TOT > ATE) and disadvantageous selection into rebate (TOT ≤ ATE). Exploratory evidence links rebate take-up more to consumption motives than to altruism.","tokens_in":30701,"tokens_out":1287,"duration_ms":9845,"significance":"If the result holds, it identifies a distinct implementation margin—voluntary take-up—that can reverse the policy ranking of price-equivalent instruments even after budget constraints and comprehension are controlled. The design is carefully pre-registered, uses a nationwide quota-balanced sample, equalizes feasible sets, and reports ITT/LATE/TOT/TOU estimates with multiple robustness checks (covariate OLS, Tobit, extensive margin, 10\times hypothetical endowment, Romano–Wolf and List–Shaikh–Xu corrections). These features make the paper a useful contribution both to the rebate-versus-matching literature and to the broader literature on selection into social programs and the scalability of experimental results. The finding that matching preserves effectiveness under opt-in while rebate does not is policy-relevant for tax-based charitable subsidies.","major_comments":[{"comment":"Section 6 and the Eckel–Grossman (2017) naturalistic acceptance gap (rebate 38% vs. matching 73%) already flag the external-validity limit of the symmetric low-cost default-off click. The central claim that self-selection reshapes the ranking is therefore best read as an experimental demonstration of an implementation margin, not as a direct prediction for real-world tax rebates (post-donation claims, filing costs, timing after the gift is chosen). The manuscript should state more sharply in the abstract and introduction that the ranking result is conditional on this standardized take-up technology, and should treat differential procedural burden as a complementary rather than competing force.","section":null},{"comment":"Section 5.1 and Figure 2: the TOU for rebate is recovered by dividing (ATE − ITT) by the non-taker share. This is valid under the maintained assumption that the compulsory ATE is the population average of TOT and TOU, but it is sensitive to any difference in the composition of the compulsory and opt-in samples beyond take-up. The paper should report a formal test of equality of baseline covariates between the compulsory and opt-in arms within each scheme (beyond the overall balance in Table 1) and discuss whether any residual imbalance could reverse the TOT ≤ ATE ordering for rebate.","section":null},{"comment":"Appendix D.4 and Section 4.1: the high-comprehension restriction (six or more of eight correct) is not pre-specified, yet it is the focal sample for the main ranking claim (Table 2 cols 5–6; Table 3). The stepwise presentation is transparent, but the paper should also report the opt-in ITT gap and LATE ordering on the full pre-registered sample as a primary robustness check in the main text, not only the low-comprehension columns, so that readers can judge how much the ranking depends on the post-hoc cut.","section":null}],"minor_comments":[{"comment":"Table 3 and Figure 2: report the exact p-value for the TOT-versus-ATE contrast under rebate (currently only the matching contrast is given as p=0.003) so that the strength of the disadvantageous-selection claim is transparent.","section":null},{"comment":"Section 4.2: the Poisson specification of Chen and Roth (2024) is a useful complement, but the mapping from estimated elasticities to γ should note that the 95% CI for γ includes both zero and values near the literature average; the text currently emphasizes consistency with γ=0 more strongly than the width of the interval warrants.","section":null},{"comment":"Appendix Table C.1 and Section 3.4: the dual-upper-limit design is clear, but a short sentence in the main text reminding readers that non-takers in the opt-in arms are analyzed under the control upper limit would reduce the chance of misreading the budget-equalization procedure.","section":null},{"comment":"Figure 1 is referenced but the caption is minimal; a one-sentence description of the 2\times2 plus control structure would help readers who encounter the figure before the text.","section":null},{"comment":"Minor typographical issues: occasional missing spaces after commas in the abstract and introduction, and inconsistent hyphenation of “opt-in” versus “opt in” as a verb.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid experimental contribution that cleanly isolates an under-studied implementation margin. The external-validity caveat is already acknowledged by the authors and does not undermine the internal claim. I would not require a field follow-up for acceptance; a clearer framing of the standardized take-up technology and the full-sample robustness check would be sufficient. Fit for a general-interest applied micro or public-finance journal is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the opt-in margin. After they equalize budgets and restrict to high-comprehension subjects, compulsory 1:1 matching and 50% rebate look the same on total amount received (26 JPY gap, p=0.351). Once people choose whether to take the subsidy, the matching advantage reappears at ~98 JPY—nearly four times larger—because rebate ITT collapses while matching ITT stays near its ATE. LATE ordering is advantageous selection into matching (TOT>ATE) and the reverse for rebate. That re-ranking under voluntary take-up is not in the prior compulsory literature (Eckel-Grossman, Blumenthal, Higgs-Uler, etc.).\n\nDesign is careful: pre-registered 2×2, quota-balanced national sample of 2,400, dual upper-limit budget fix, comprehension checks, Fowlie-style TOT/TOU, and sensible robustness (covariate OLS, Tobit, extensive margin, 10× hypothetical, Romano-Wolf). Theory appendix maps cleanly onto Hungerman-Ottoni-Wilhelm; structural e and γ are estimated after the fact for interpretation only and do not force the ranking. Citations are honest about what was already known.\n\nSoft spots are real but secondary. The high-comprehension cut and the consumption-motive exploration were not fully pre-specified (they disclose this). No public data/code yet. The binding limit is external validity: a symmetric low-cost default-off click is not tax filing or a matching-campaign channel, and they already flag the Eckel-Grossman 2017 acceptance gap. That does not break the experimental claim; it limits how hard you can push the policy ranking.\n\nThis is for people who work on charitable subsidies, tax incentives, or selection into social programs. The reduced-form result is sharp enough that a serious editor should send it out. I would cite the opt-in re-ranking and bring it to reading group.","headline":"Clean 2×2 experiment shows voluntary take-up restores a large matching advantage via opposite selection patterns; the internal result is solid, the external-validity claim is the soft spot.","tokens_in":31308,"tokens_out":508,"would_cite":true,"duration_ms":5779,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When donors must opt in, matching stays effective while rebate loses half its power—even though both give the same price under forced assignment.","keywords":["charitable subsidy","tax incentive","matching","rebate","policy implementation","self-selection","scaling","LATE"],"falsifier":"A field comparison of matching and rebate take-up and giving in which real administrative costs and timing are allowed to differ by instrument (for example, post-donation tax filing versus designated matching channels) and the ranking of total amounts received is measured under those natural frictions.","tokens_in":31300,"feed_emoji":"🎁","tokens_out":656,"duration_ms":6777,"temperature":0.7,"pith_summary":"Theory says a 1:1 match and a 50% rebate should produce the same charitable giving because they impose the same out-of-pocket price. Earlier lab results favored matching, but later refinements that equalize budgets and fix confusion largely erase that gap when people are assigned to a scheme. This paper asks what happens when take-up is voluntary, the usual real-world case. In a nationwide incentivized experiment with 2,400 Japanese adults, the authors cross subsidy type with compulsory versus opt-in assignment. Once budgets are equalized and only high-comprehension participants are kept, total amounts received by the charity are statistically indistinguishable under compulsory assignment. Under opt-in the matching advantage reappears and nearly quadruples, not because matching improves, but because rebate loses roughly half its effectiveness. Local average treatment effect estimates show advantageous selection into matching and disadvantageous selection into rebate. The policy ranking of two price-equivalent instruments can therefore be reversed by who chooses to use them.","feed_headline":"Opt-in nearly quadruples the matching edge over rebate","feed_subtitle":"Same price under forced assignment; under voluntary take-up, rebate loses half its effect","key_machinery":"A 2×2 design that crosses subsidy type (1:1 matching vs 50% rebate) with assignment rule (compulsory vs opt-in), combined with budget-equalized feasible sets and a LATE decomposition into treatment-on-the-treated and treatment-on-the-untreated, which isolates how self-selection changes realized effectiveness.","core_discovery":"After equalizing budget constraints and restricting to high-comprehension participants, total giving is statistically indistinguishable under compulsory 1:1 matching versus 50% rebate. Under voluntary opt-in the matching advantage re-emerges and nearly quadruples because the rebate intent-to-treat effect falls by about half while matching stays near its compulsory average treatment effect, with LATE patterns of advantageous selection into matching and disadvantageous selection into rebate.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Opt-in nearly quadruples matching edge as rebate loses half effect","Voluntary take-up nearly quadruples matching advantage over rebate","Matching edge nearly quadruples under opt-in as rebate weakens","Self-selection nearly quadruples matching lead over rebate","Rebate effect halves under opt-in nearly quadrupling match edge"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The experiment’s symmetric, low-cost default-off click is assumed to isolate the same selection-on-gains margin that governs real-world tax rebates and matching campaigns, where procedural burdens and timing differ sharply by instrument.","fun_headline_variants_meta":{"raw":{"variants":["Opt-in nearly quadruples matching edge as rebate loses half effect","Voluntary take-up nearly quadruples matching advantage over rebate","Matching edge nearly quadruples under opt-in as rebate weakens","Self-selection nearly quadruples matching lead over rebate","Rebate effect halves under opt-in nearly quadrupling match edge"]},"model":"grok-4.5","effort":"low","cost_usd":0.00817,"raw_usage":{"total_tokens":1809,"prompt_tokens":671,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":81700000,"prompt_tokens_details":{"text_tokens":671,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1070,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":671,"tokens_out":68,"duration_ms":8380,"temperature":1.0,"reasoning_tokens":1070,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:48:06.190987+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A field comparison of matching and rebate take-up and giving in which real administrative costs and timing are allowed to differ by instrument (for example, post-donation tax filing versus designated matching channels) and the ranking of total amounts received is measured under those natural frictions.","supporting_citations":[],"review_version":1}