{"id":"75e4f5e7-d053-42aa-8da6-2d4e24a6e86a","arxiv_id":"2506.01385","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Using survey data from 159,000 users and a regional input-output model, the study estimates Taipei's voucher program has a GDP multiplier of 0.97 without behavioral responses and up to 1.76 when substitution and induced consumption are included, with accommodation vouchers most effective.","lead":"This paper evaluates Taipei's 2022 digital consumption voucher program using a survey of about 159,000 users and a regional input-output model. It finds that accommodation vouchers trigger the most new spending and that the program's GDP multiplier rises from 0.97 to as high as 1.76 when consumers' substitution and extra-spending behavior is included.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 5 is not reproducible from Tables 2 and 3: the headline multipliers (e.g., pessimistic sports 0.79 vs 0.61 implied by (1−ES)(1+IC)) appear to rest on unexplained aggregation, so the central 1.762 estimate lacks internal support.","rationale":"I agree with the reader that the self-reported counterfactual questions and the Section 4.4 bias correction are a serious identification problem, and that a conditional verdict is right. However, examining Table 5 against Tables 2 and 3 surfaces a sharper, purely internal problem: the multiplier inputs do not follow from the stated formula. The discrepancies are too large to be rounding (sports pessimistic: 0.614 implied versus 0.79 reported; accommodation pessimistic: 1.311 implied versus 1.58 reported). Since Table 6's 1.762 is a mechanical function of Table 5, the headline claim is not currently reproducible from the published estimates. This does not require assuming the authors are wrong; it requires an explanation of the aggregation or a corrected table. The reader's conditional verdict and request for validation (transaction-level data or a control group) stand; I would not escalate to rejection because the arithmetic and methodological concerns are potentially fixable within the existing study design.","tokens_in":20471,"tokens_out":16571,"duration_ms":175890,"concrete_test":"Recompute Table 5's multipliers from the reported overall lower/upper bounds in Tables 2 and 3 using exactly (1−ES_k)(1+IC_k), then re-run the IO calculation in Table 6 with the corrected induced-demand values. In particular, compute ΔF_pess = Original × (1−ES_upper)(1+IC_lower) and ΔF_opt = Original × (1−ES_lower)(1+IC_upper) for each voucher type, and compare the resulting total multiplier with the reported 1.143 and 1.762. If the optimistic multiplier changes by more than about 0.1 or falls below 1.5, the paper must either supply the omitted aggregation code or revise the headline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — the output multiplier rises from 0.969 to 1.762 — rests directly on Table 5's 'behaviorally adjusted' demand figures, which are then fed into the regional IO model (Section 5.6). But Table 5 does not reproduce from the paper's own reported behavioral bounds. Section 5.6 states the adjustment is (1−ES_k)×(1+IC_k). Using the overall bounds from Tables 2 and 3: accommodation pessimistic gives (1−0.24)(1+0.725)=1.311, yet Table 5 reports 1.58; sports pessimistic gives (1−0.728)(1+1.259)=0.614, yet Table 5 reports 0.79; sports optimistic gives (1−0.405)(1+1.896)=1.723, yet Table 5 reports 1.63. These gaps are far beyond rounding (approximately 29% relative error in the sports pessimistic case). Table 6's optimistic multiplier 1.762 inherits these inputs, so unless the omitted aggregation is explained (e.g., bootstrap percentile of the product, subgroup-weighted means, or a different choice of bounds), the headline estimate is not currently supported by the paper's own tables. This is more immediately load-bearing than the also-valid concern about self-reported counterfactuals: even granting the survey answers, the published arithmetic does not close.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the Taipei Bear Vouchers 2.0 program using a large, platform-administered survey of 159,211 users and a regional input–output model built from Taiwan's national IO table via the simple location quotient method. It estimates expenditure substitution rates and induced consumption rates for six voucher types, applies a bootstrap-based bias correction that subtracts the minimum subgroup mean, and combines the two behavioral parameters into a final-demand adjustment factor (1−ES_k)(1+IC_k). The headline result is that the program's output multiplier rises from 0.969 in a behavioral baseline to 1.762 in an optimistic scenario, with accommodation vouchers showing the strongest amplification. The paper also examines treatment intensity from an additional round of bonus vouchers.","tokens_in":20865,"tokens_out":7793,"duration_ms":80882,"significance":"If the estimates are correct, the paper offers a useful policy case study: it uses an unusually large survey of verified voucher users, distinguishes six voucher categories, attempts an explicit bias-correction design, and constructs a regional IO table with transparent methodology. The finding that voucher effectiveness varies widely by sector—sports vouchers largely substitute for planned spending while accommodation vouchers induce substantial new spending—is practically relevant for designing targeted stimulus. However, the headline multiplier is currently undermined by an internal arithmetic inconsistency: Table 5 does not reproduce from Tables 2 and 3 using the formula stated in Section 5.6, and the survey-based counterfactual parameters are not validated against transaction-level data. The contribution is therefore conditional on a successful reanalysis of the central calculation.","major_comments":[{"comment":"The behaviorally adjusted multipliers in Table 5 do not reproduce from Tables 2 and 3 using the formula (1−ES_k)(1+IC_k) stated in Section 5.6. For example, the accommodation pessimistic adjustment is (1−0.24)(1+0.725)=1.31, yet Table 5 reports 1.58; the sports pessimistic adjustment is (1−0.728)(1+1.259)=0.61, yet Table 5 reports 0.79; and the sports optimistic adjustment is (1−0.405)(1+1.896)=1.72, yet Table 5 reports 1.63. These gaps are far beyond rounding, especially the 29% relative error in the sports pessimistic case. Because the optimistic output multiplier 1.762 in Table 6 is computed from these adjusted demand inputs, the headline result is not currently supported by the paper's own tables. The authors should either document the exact aggregation rule (subgroup weighting, bootstrap percentile, or a different choice of bounds) or recompute Tables 5 and 6.","section":"5.6, Table 5"},{"comment":"The bias correction treats bB_k = min_j \\hat y_{jk} as an upper bound on reporting bias for every subgroup. This is an identifying assumption, not a result. Tables 2 and 3 show large and systematic subgroup variation—for instance, monotone age gradients in induced consumption and large residence differences—so subtracting the minimum subgroup mean is likely to overcorrect subgroups whose true effects are genuinely low. This directly affects the lower-bound estimates and the pessimistic multiplier. The paper should provide external validation against transaction-level data from TaipeiPASS or a sensitivity analysis with alternative bias bounds before the pessimistic scenario can be taken as credible.","section":"4.4, Assumptions 2–3 and Equations (6)–(7)"},{"comment":"Equation (8) defines y = (I−A)^{-1}(ΔF)∘VA and calls y 'output,' but multiplying the Leontief solution by value-added coefficients makes y a value-added (GDP) vector, not an output vector. The paper then defines the output multiplier as the change in GDP divided by the original policy expenditure, which compounds the terminological confusion. The distinction between output and value added should be made precise throughout Section 5.6, and the baseline GDP figure of NT$566.32 million should be reconciled with the final-demand input of NT$584.53 million and the SLQ-based regional leakage assumptions.","section":"4.5, Equation (8), Table 6"},{"comment":"Equation (5) defines IT_k as a per-respondent average difference in additional out-of-pocket spending, but Table 4 reports values labeled 'NT$ millions' ranging from 28.55 to 860.70. If these are aggregate program effects, Equation (5) omits the relevant sample sizes; if they are per-person amounts, the unit label is incorrect. Either way, the treatment-intensity results cannot be interpreted without resolving this discrepancy, and the reported magnitudes should be checked against the program budget and voucher face values.","section":"5.4, Table 4, Equation (5)"},{"comment":"The identification of the two key behavioral parameters rests on self-reported counterfactual questions—whether the purchase 'would have occurred' without the voucher and how much was spent beyond the face value. The paper calls the data 'verified user-level survey data,' but verification appears to mean only that respondents were actual voucher users, not that their counterfactual reports were checked against transaction records. Since TaipeiPASS tracks redemptions, the authors should attempt a validation exercise or clearly state the limitation in the interpretation of the multipliers.","section":"3, 5.2, 5.3"}],"minor_comments":[{"comment":"The caption of Figure 2 reads 'Expenditure Substitution Rate,' but the figure displays induced consumption effects; the caption should be corrected.","section":"Figure 2"},{"comment":"The abstract and text state 159,211 valid responses, while Table 1's column total sums to 159,221; the discrepancy should be fixed.","section":"Table 1 and abstract"},{"comment":"Section 5.5 refers to 'Taiwan's 2009 paper-based voucher program,' but Section 2.2 and the cited Kan et al. (2017) study describe the 2008 program; the year should be consistent throughout.","section":"5.5"},{"comment":"Equation (5) is introduced as an intensity-of-treatment index, but the text ends the paragraph with 'it delivers different meanings compared with the induced income rate'; 'induced income rate' is an undefined term and appears to be a typo.","section":"4.3"},{"comment":"The input–output coefficient matrix in Table 7 is difficult to read because the column alignment is compressed; a cleaner presentation with sector abbreviations would help reproducibility.","section":"Appendix, Table 7"},{"comment":"In several places the paper uses 'lower-est' and 'upper-est' without defining the terms in one place; define them once in Section 4.4 and use them consistently.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The Table 5 inconsistency is the main barrier to publication: the central multiplier claim is not currently supported by the paper's own tables. If the authors can provide the exact aggregation used and recompute the IO results, the paper could become publishable after a major revision. I would also want the editor to confirm that the proprietary survey data access is secure and that the authors' institutional arrangement with Taipei City Government does not create an undisclosed conflict. The paper's scope fits an applied economics or regional economics journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe first thing to know: this paper has a genuinely useful dataset and asks the right question, but its headline result—that the Taipei Bear Vouchers 2.0 output multiplier rises from 0.969 to 1.762 once behavior is accounted for—is not supported by the paper's own numbers. Table 5 does not reproduce from the formula stated in Section 5.6. Using the overall bounds from Tables 2 and 3, the pessimistic accommodation multiplier should be about 1.31, not 1.58; the pessimistic sports multiplier should be about 0.61, not 0.79. The optimistic sports figure is also off. These are not rounding issues. Since Table 6's headline multipliers are fed by Table 5, the 1.762 estimate is currently on weak internal footing.\n\nWhat the paper does well: it draws on a large verified survey (159,211 responses) from the TaipeiPASS platform, clearly defines expenditure substitution and induced consumption, and gives a useful breakdown across six voucher types. The added value relative to Kan et al. (2017) is real, and the stratified-minimum bias-correction is a reasonable, transparent way to bound self-reporting error. The qualitative findings—accommodation vouchers being strongly incremental, sports vouchers largely substituting for planned spending—are plausible and consistent with the earlier literature.\n\nThe soft spots, in order of importance. First, the internal inconsistency above is load-bearing; the paper needs to explain the aggregation or correct the table. Second, the behavioral parameters themselves come from self-reported counterfactual answers, so even a clean multiplicative adjustment inherits that subjectivity. No control group, and the bias-correction assumes the smallest subgroup mean is an upper bound on bias, which is a heuristic. Third, Table 4's units appear off: values in 'NT$ millions' are not consistent with Equation (5), which produces per-capita amounts.\n\nThe stress-test note holds up. The central claim needs either a solid explanation or a revised estimate. That said, this is fixable. The data and research design are worth a serious referee, so I'd send it to review with a request that the referee verify the arithmetic. Not a desk reject; a 'please fix the numbers' revise-and-resubmit.\n\nFor your reading group, it's a good example of why you should always back out the main table from the stated formula.","headline":"Useful data and a real question, but the headline multiplier does not reproduce from the paper's own tables.","tokens_in":21309,"tokens_out":6410,"would_cite":false,"duration_ms":59658,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Consumer behavior—how much voucher spending replaces planned purchases and how much spills into extra out-of-pocket spending—can double a voucher program's measured economic impact, raising Taipei Bear Vouchers 2.0's output multiplier…","keywords":["expenditure substitution","induced consumption","consumption vouchers","fiscal stimulus","regional input-output model","output multiplier","self-reporting bias","COVID-19 recovery"],"falsifier":"Compare the survey's self-reported substitution and induced-spending answers against actual transaction-level records from the TaipeiPASS redemption platform; if the true substitution rate equals the all-substitution baseline scenario, the multiplier would be 0.969, below the paper's behavioral range.","tokens_in":20278,"feed_emoji":"💳","tokens_out":4767,"duration_ms":49239,"temperature":0.7,"pith_summary":"This paper evaluates Taipei Bear Vouchers 2.0, a digital consumption voucher program, using 159,211 verified survey responses from actual users. It aims to measure how much voucher spending substitutes for planned purchases (expenditure substitution) versus generates new out-of-pocket spending (induced consumption), and to feed those behaviorally adjusted amounts into a regional input–output model. Doing so raises the program's output multiplier from 0.97, when behavior is ignored, to as high as 1.76, while showing that accommodation vouchers—low substitution, high induced spending—are the most effective type and sports vouchers often replace existing spending. A sympathetic reader should care because the result says that the design and targeting of vouchers, not just their face value, determines whether a fiscal stimulus program delivers a multiplier above one.","feed_headline":"How user behavior doubles Taipei's voucher multiplier to 1.76","feed_subtitle":"Self-reported substitution and extra spending turn a 0.97 baseline into a 1.76 output gain.","key_machinery":"The central identity is the behaviorally adjusted final demand per voucher, $(1-ES_k)\\times(1+IC_k)$, where $ES_k$ is the self-reported substitution rate and $IC_k$ is the induced consumption rate computed from survey bracket midpoints. This adjusted demand enters a 19-sector Taipei regional input–output model $y=(I-A)^{-1}(\\Delta F)\\circ VA$, with the regional coefficient matrix $A$ built by Simple Location Quotient adjustment from Taiwan's 2016 national input–output table. Uncertainty is handled by a stratified bootstrap that produces optimistic (no bias correction) and pessimistic (subtracting the minimum subgroup mean) bounds on the estimates.","core_discovery":"The paper claims that for the Taipei Bear Vouchers 2.0 program, consumer behavioral responses are first-order for evaluating fiscal stimulus: sports vouchers substitute for 40.5% to 72.8% of planned spending, while accommodation vouchers substitute for only 12.0% to 24.0%; induced consumption is highest for accommodation at 72.5% to 251.6% of the voucher's face value. Applying the adjustment factor $(1-ES_k)\\times(1+IC_k)$ to final demand and running the regional Leontief inverse gives a GDP impact that rises from NT\\$566 million in the baseline to NT\\$1,030 million in the optimistic scenario, with an output multiplier of 1.762 versus 0.969. The authors argue that ignoring these behavioral parameters understates the program's economic contribution and that untargeted sectors receive substantial indirect gains through inter-industry linkages.","pith_inferences":["The same behavioral-adjustment formula could be applied to other voucher programs, but the estimated substitution and induced-consumption parameters are specific to Taipei's demographics, urban density, and program rules; transferability to rural or national programs is untested.","If administrator transaction records ever become linkable to individual users, the self-reported counterfactual answers could be validated; a finding that users systematically overstate induced spending would collapse the optimistic multiplier toward the pessimistic bound.","The regional model's Simple Location Quotient assumption may not hold for a small service-oriented city, because real supply chains could leak more demand to other regions, making the true city-level multiplier lower than 1.762.","The 'unexpected policy' intensity result suggests an optimization margin: governments can raise stimulus per dollar by surprising consumers rather than pre-announcing top-ups, though repeated surprise rounds could lose their effect."],"forward_implications":["Program design should steer vouchers toward categories with low substitution and high induced consumption, because accommodation-like categories maximize incremental output.","Ignoring consumer behavior understates the multiplier by roughly half in the optimistic scenario (0.969 to 1.762), so cost–benefit analyses of voucher stimuli should measure these behavioral parameters.","Raising voucher face value can produce meaningful marginal spending, especially when the increase is unexpected, suggesting that surprise top-up rounds are a cost-effective form of stimulus.","Indirect gains in untargeted sectors, such as professional services, mean the program's benefit extends beyond the directly targeted retail, food, and accommodation industries.","Digital vouchers with sector restrictions can outperform cash transfers, which typically show marginal propensities to consume below 0.4–0.6, because they channel spending to high-multiplier local sectors."],"supporting_citations":[{"why":"Provides the baseline paper-voucher MPC of 0.243, the comparison showing that restricted digital vouchers outperform unrestricted paper vouchers.","marker":"Kan et al. (2017)"},{"why":"Digital coupon evidence from Shaoxing, giving the per-yuan spending benchmark (3.07) and the method of measuring induced consumption that this paper contrasts.","marker":"Xing et al. (2023)"},{"why":"Ningbo digital coupon study, providing the revenue per yuan and welfare results that this paper compares against.","marker":"Chen et al. (2025b)"},{"why":"Methodological basis for applying a regional input–output model to evaluate Taiwan's voucher programs.","marker":"Hua et al. (2022)"},{"why":"Semi-closed input–output model for short-run fiscal stimulus effects, guiding the estimation of output multipliers.","marker":"Chen et al. (2016)"},{"why":"Standard reference for the Leontief inverse and regional input–output construction, underpinning the multiplier calculation.","marker":"Miller and Blair (2009)"},{"why":"Origin of the Simple Location Quotient regionalization method used to build the Taipei input–output table.","marker":"Isard and Kuenne (1953)"},{"why":"Official implementation report of the predecessor program, the source of comparison substitution and induced-consumption rates.","marker":"Taipei City Government (2022)"}],"fun_headline_variants":["Behavior doubles Taipei voucher multiplier to 1.76","Accommodation vouchers outdo sports in Taipei payout","Voucher type and consumer behavior drive multiplier to 1.76","Low substitution boosts accommodation voucher impact in Taipei","From 0.97 to 1.76: Taipei vouchers rely on user behavior"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results depend on voucher users' self-reports of whether they would have made the purchase anyway and how much extra they spent; the paper's bias correction assumes the lowest-reported subgroup mean is an upper bound on any over-reporting, so if users systematically overstate new spending, the corrected multipliers are too high.","fun_headline_variants_meta":{"raw":{"variants":["Behavior doubles Taipei voucher multiplier to 1.76","Accommodation vouchers outdo sports in Taipei payout","Voucher type and consumer behavior drive multiplier to 1.76","Low substitution boosts accommodation voucher impact in Taipei","From 0.97 to 1.76: Taipei vouchers rely on user behavior"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000311,"raw_usage":{"total_tokens":1744,"prompt_tokens":888,"completion_tokens":856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":769}},"tokens_in":504,"tokens_out":856,"duration_ms":9194,"temperature":1.0,"reasoning_tokens":769,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:43:26.335218+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the survey's self-reported substitution and induced-spending answers against actual transaction-level records from the TaipeiPASS redemption platform; if the true substitution rate equals the all-substitution baseline scenario, the multiplier would be 0.969, below the paper's behavioral range.","supporting_citations":[],"review_version":1}