{"id":"16ee0d66-f63b-48b6-a98a-2e4b5e01bcef","arxiv_id":"2412.04265","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In multi-cutoff regression discontinuity designs, constant-bias extrapolation is unreliable when the running variable is manipulable, and under monotonicity plus dominance the extrapolated treatment effect is sharply bounded by two estimable regression functions.","lead":"This paper studies whether treatment effects estimated at a cutoff in multiple-cutoff regression discontinuity designs can be extrapolated to other values of the running variable. It shows the standard \"constant bias\" assumption can fail when the running variable is partially manipulable, and derives bounds on extrapolated effects under monotonicity and dominance assumptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 2's 'if and only if' is not proven: the only-if step needs supp(ϵ) connected and e*(ϵ) continuous/surjective to turn pointwise density equalities into periodicity on [l-b,h-a], and Assumption 2.3 omits this.","rationale":"The reader's weakest-assumption field points to monotonicity of µ0,l, which is indeed untestable and acknowledged in Remark 8; I do not view that as the most load-bearing issue because untestability is normal for shape restrictions and the paper is transparent about it. The Proposition 2 proof gap is a correctness issue in a headline theoretical result, and the paper itself flags no missing condition. I therefore focus on it. If the iff is weakened to a one-way statement ('periodic density implies cutoff-independent effort') plus the numerical examples, the qualitative lesson that partial manipulability can break constant bias survives, but the strong characterization claimed in Proposition 2 does not. This is a repair rather than a rejection, so the conditional verdict stands. The alternative bounds and their sharpness proof appear internally sound, and the empirical illustrations are reasonable; the main other request is to make code auditable.","tokens_in":36154,"tokens_out":11561,"duration_ms":125202,"concrete_test":"Fix l=0, h=1, u(s)=s, s(e)=e, y(e)=e, K(e,ϵ)=0.5(2-ϵ)e^2, and let ϵ have support {0,1}. Choose a smooth non-periodic density fηs (e.g., a mixture of two normals) and tune its parameters so that the first-order conditions give fηs(-e*(0))=fηs(1-e*(0)) and fηs(-e*(1))=fηs(1-e*(1)) while fηs is not 1-periodic on [-b,1-a]. Solving the FOCs for both cutoffs would show e*_l=e*_h for all realizations, contradicting the claimed 'only if.' A purely analytic alternative: add 'supp(ϵ) is an interval and e* is interior' to Assumption 2.3 and check whether the existing proof goes through unchanged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3.2's Proposition 2 is the theoretical basis for the paper's negative message that partial manipulability breaks the constant-bias assumption. The proof in Appendix A.1 shows that if e*_l(ϵ)=e*_h(ϵ)=e*(ϵ), then fηs(l-s(e*(ϵ)))=fηs(h-s(e*(ϵ))) for every ϵ. It then concludes that fηs is periodic with period h-l on [l-b,h-a], where a=inf_ϵ s(e*(ϵ)) and b=sup_ϵ s(e*(ϵ)). This inference requires the set {s(e*(ϵ)): ϵ∈supp(ϵ)} to be the entire interval [a,b], so that every z in [l-b,l-a] is of the form l-s(e*(ϵ)). Continuity of e* gives this only if supp(ϵ) is connected and the solution is interior for all ϵ. Assumption 2.3(v) only states that the support of ϵ is identical across groups; it does not require connectedness, and the FOC is assumed only through strict concavity, not an interior solution with nonzero second derivative. For a disconnected support, e.g., two ability types, the equality pins down f at two pairs of points, far from full periodicity. Hence the stated iff is stronger than what is established. Because the bound results in Section 3 do not rely on Proposition 2, the paper's main identification contribution survives, but the microfoundation claim requires either an added support/regularity condition or a weakened conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies extrapolation of treatment effects away from the cutoff in multi-cutoff sharp regression discontinuity designs. It first builds a micro-founded decision model in which agents choose effort that affects a test score and a future outcome; the model is used to argue that the constant-bias assumption of Cattaneo et al. (2021) is plausible when the running variable is non-manipulable and cutoff assignment is as-if random (Proposition 1), but that it can fail when the running variable is partially manipulable (Proposition 2 and Examples 1-2). The paper then proposes a complementary partial-identification strategy: under continuity, monotonicity of the control outcome functions mu_{0,c}, and dominance mu_{0,l} <= mu_{0,h} on (l,h), the extrapolated effect tau_l(xbar) is pointwise sharply bounded by mu_{1,l}(xbar)-mu_{0,h}(xbar) and mu_{1,l}(xbar)-mu_{0,l}(l) (Theorem 1). The paper develops local-linear estimation with robust bias correction, pointwise confidence intervals, and a multiplier-bootstrap uniform confidence band, and illustrates the methods with the SPP and ACCES scholarship programs. The online appendix contains proofs, a fuzzy-RD extension, and simulations.","tokens_in":36466,"tokens_out":16345,"duration_ms":177003,"significance":"Conditional on the maintained assumptions, Theorem 1 is a clean and practically useful result: the bounds are simple, pointwise sharp, invariant to increasing monotone transformations of the outcome, and the paper honestly discloses that they are not uniformly sharp. The uniform inference procedure is nontrivial and is supported by a detailed asymptotic argument in the online appendix, simulations, and replication code, all of which are valuable. The identification assumptions are stated ex ante rather than fitted to the data used to evaluate the claims, and the numerical examples are clearly illustrative rather than calibrated. The main weakness is Proposition 2: its 'if and only if' is not proved under the stated assumptions, and because that proposition underpins the paper's negative message about partial manipulability, the manuscript needs a substantive revision before that part of the contribution is convincing. The Section 3 identification contribution does not depend on Proposition 2 and therefore survives the revision.","major_comments":[{"comment":"The proof of Proposition 2 does not establish the stated 'if and only if' under Assumption 2.3. In the only-if direction, the first-order condition gives f_eta^s(l - s(e*(epsilon_i))) = f_eta^s(h - s(e*(epsilon_i))) for each epsilon_i, and the proof then asserts that f_eta^s is periodic on [l-b, h-a], where a = inf_epsilon s(e*(epsilon)) and b = sup_epsilon s(e*(epsilon)). This inference requires the set {s(e*(epsilon)) : epsilon in supp(epsilon)} to be the entire interval [a,b], which needs supp(epsilon) to be connected and e* to be continuous on the support. Assumption 2.3(v) only states that the support of epsilon is identical across groups; it does not require connectedness, and continuity of e* is not guaranteed by the stated strict concavity and differentiability conditions alone (the implicit function theorem needs a nonzero second derivative at the optimum). With a disconnected support, for example two ability types, the equality pins down f_eta^s only at finitely many pairs of points, far short of periodicity on the whole interval. The converse direction is also incomplete: the proof claims that under periodicity the probability identity holds 'for any e_i', but the assumed periodic interval is defined through the optimal efforts. For a feasible effort with s(e_i) outside [a,b], the endpoints l - s(e_i) and h - s(e_i) lie outside the periodic interval, so the quantity Q is not shown to be constant in e_i. Thus the equivalence of the two decision problems is not proved without additional assumptions on the range of s over the whole choice set. Because Proposition 2 is the theoretical basis for the claim that partial manipulability generically breaks the constant-bias assumption, the authors should either add explicit support-connectedness, interiority, and range conditions and prove the required continuity, or restate Proposition 2 as a weaker necessary condition plus a separate sufficiency result.","section":"Section 2.3.2; Appendix A.1"},{"comment":"Assumption 3.1 requires mu_{0,l} to be weakly increasing on (l,h), but mu_{0,l} is not observed for x > l in a sharp RD design because the treated outcome is the one observed there; Remark 8 honestly admits this untestability. This assumption is load-bearing for the upper bound in Theorem 1, because the bound mu_{1,l}(xbar) - mu_{0,l}(l) is valid only if mu_{0,l}(xbar) >= mu_{0,l}(l). The empirical summaries in Sections 4.1.3 and 4.2.3, however, draw substantive policy conclusions from the full bounds (for example, that large negative effects are ruled out). To make the reported message match the identifying power actually available, the paper should either prominently qualify that the upper bound depends on the untestable monotonicity of mu_{0,l}, or report a lower-bound-only analysis that relies only on Assumption 3.2 and is therefore robust to violations of that monotonicity.","section":"Section 3.1.1; Remark 8; Section 4.2.3"}],"minor_comments":[{"comment":"The displayed confidence band uses bV_L^{-1/2}(x) and bV_U^{-1/2}(x), while Step 3 and the proof in Online Appendix S1.1.1 use bV_L^{1/2}(x) and bV_U^{1/2}(x); the notation should be made consistent.","section":"Section 3.2.3, Step 4"},{"comment":"In the first-order condition in the proof of Proposition 2, the first term appears to have a missing parenthesis: it should be u'(s(e_i^*)) s'(e_i^*).","section":"Appendix A.1"},{"comment":"The statement of Proposition 2 uses inf_epsilon s(e*(epsilon)) and sup_epsilon s(e*(epsilon)) without first defining a and b and without specifying that the infimum and supremum are taken over the support of epsilon; this should be made explicit.","section":"Proposition 2"},{"comment":"The demonstration that the bounds are not uniformly sharp is informal and based on a figure; a formal counterexample with explicit functions would make the claim easier to verify.","section":"Online Appendix S1.2"},{"comment":"The invariance claim should specify that the identification conditions are invariant to increasing monotone transformations of the outcome; for decreasing transformations the monotonicity and dominance directions reverse, and Corollary 1 applies instead.","section":"Remark 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well organized and the main partial-identification result in Section 3 is likely publishable after the decision-model section is corrected. The stress-test concern about Proposition 2 lands: the proof gap is real and affects the paper's negative message about partial manipulability, but it is fixable by adding assumptions or weakening the claim. No citation or scope concerns; the contribution relative to Cattaneo et al. (2021) and Sun (2023) is clear."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the main contribution is solid: under monotonicity of the control outcome and dominance across cutoffs, the paper derives pointwise sharp bounds on extrapolated treatment effects, and backs them with a uniform inference procedure. That is a clean, implementable result that should be useful to applied researchers. The authors also honestly disclose that the bounds are not uniformly sharp, which is the right call. Second, the microfoundation analysis of the constant bias assumption is directionally interesting but the headline 'if and only if' in Proposition 2 is stronger than the proof supports.\n\nThe proof of Proposition 2 shows that if optimal effort is cutoff-invariant, then fηs(l − s(e*(ϵ))) = fηs(h − s(e*(ϵ))) for each ability type ϵ. Concluding that fηs is periodic on the whole interval [l−b, h−a] requires that the set {s(e*(ϵ))} covers [a, b]. That needs ϵ to have connected support and the solution to be interior with enough regularity; Assumption 2.3 only says support is identical across groups. With a two-point ability distribution you get equality at two pairs of points, not periodicity. So the only-if direction fails as stated. This is not fatal to the paper's main identification theorem, which does not use Proposition 2, but the claim needs to be weakened to a sufficient condition or given extra assumptions.\n\nThe bounds themselves are on firmer ground. The lower bound uses dominance, whose endpoint version is testable; the upper bound uses monotonicity of the lower-cutoff control function on (l, h), which is untestable. The paper acknowledges this in Remark 8, yet the upper bound is still load-bearing for the headline results, so applied readers should treat it as a substantive assumption, not a mild shape restriction. The empirical illustrations are informative: the SPP example shows tight bounds, and the ACCES example usefully shows the bounds ruling out large negative parametric extrapolations. The simulation and fuzzy RD extension are careful. The only other annoyance is that the R code is mentioned but not verifiable from the text; a commit hash would help.\n\nFor whom? Applied microeconometricians working with multi-cutoff RD and methodologists interested in partial identification will both get value. I would send this to a serious referee: the main theorem and inference procedure deserve a full review, and the referee should be asked to fix or qualify Proposition 2 and require code deposit.","headline":"A genuinely useful partial-identification result for multi-cutoff RD, with an overreaching microfoundation claim that needs a fix.","tokens_in":37003,"tokens_out":1814,"would_cite":true,"duration_ms":21402,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows when the standard constant-bias assumption for extrapolating treatment effects across multiple regression-discontinuity cutoffs fails, and offers bounds that stay valid.","keywords":["multiple-cutoff regression discontinuity","extrapolation","partial identification","constant bias assumption","monotonicity","dominance","uniform inference","decision model"],"falsifier":"Estimate the control-outcome function $\\mu_{0,l}(x)$ for $x \\in (l,h)$ using auxiliary data or a design that removes treatment for the lower-cutoff group; a non-monotone estimate invalidates the upper bound and the sharpness claim. For the decision-model failure result, estimate the density of score shocks across the two cutoff groups: an absence of $h-l$ periodicity indicates, by Proposition 2, that optimal effort responds to the cutoff and the constant-bias assumption fails.","tokens_in":35926,"feed_emoji":"📈","tokens_out":9551,"duration_ms":81300,"temperature":0.7,"pith_summary":"This paper asks when treatment effects estimated by regression discontinuity at one cutoff can be credibly extrapolated to other values of the running variable when several cutoffs exist. The authors show that the standard constant-bias assumption—that the gap between two groups' untreated outcome functions is flat—is justified only when the running variable cannot be manipulated by agents and cutoffs are assigned as if randomly. When the running variable is partially manipulable, such as a test score, a rational-agent model predicts that cutoff shifts change effort, so the assumption generically fails and extrapolations can be biased. As a complement, the paper derives bounds on the extrapolated effect for the lower-cutoff group at any point between the two cutoffs, using monotonicity of the untreated outcome functions and a dominance condition, and shows these bounds are pointwise sharp. It then develops estimation and uniform inference procedures and illustrates them on two financial-aid programs.","feed_headline":"Expect RD extrapolation bias from score manipulation; bounds fix it","feed_subtitle":"Why the standard extrapolation fails when running variables respond to effort, and how bounds restore identification.","key_machinery":"The identification machinery is the pair of bounds $\\underline{\\nabla}_l(\\bar{x}) = \\mu_{1,l}(\\bar{x}) - \\mu_{0,h}(\\bar{x})$ and $\\overline{\\nabla}_l(\\bar{x}) = \\mu_{1,l}(\\bar{x}) - \\mu_{0,l}(l)$, delivered by monotonicity of the lower-cutoff group's control outcome function (upper bound) and dominance of the higher-cutoff group's control outcome (lower bound). The theoretical machinery for the failure result is a rational-agent model in which effort $e_i$ shifts both the running variable and the control outcome; a cutoff shift changes effort whenever the idiosyncratic score-shock density $f_{\\eta^s}$ is not periodic with period $h-l$. For inference, the paper uses local linear estimation with an equivalent-kernel expansion, robust bias correction, and a multiplier bootstrap with Mammen weights to obtain uniform confidence bands.","core_discovery":"The paper's central claim is twofold. First, a microeconomic decision model shows that the constant-bias assumption—that $\\mu_{0,h}(x) - \\mu_{0,l}(x)$ is flat over $(l,h)$—is justified when the running variable is non-manipulable and cutoff assignment is as-if random, but fails generically when the running variable is partially manipulable, even when the groups are identical in ability and beliefs. Second, under continuity, monotonicity of $\\mu_{0,c}$, and dominance $\\mu_{0,l} \\le \\mu_{0,h}$, the extrapolated treatment effect for the lower-cutoff group at any $\\bar{x} \\in (l,h)$ is pointwise sharply bounded by $\\underline{\\nabla}_l(\\bar{x}) = \\mu_{1,l}(\\bar{x}) - \\mu_{0,h}(\\bar{x})$ and $\\overline{\\nabla}_l(\\bar{x}) = \\mu_{1,l}(\\bar{x}) - \\mu_{0,l}(l)$.","pith_inferences":["A practical diagnostic emerges from the decision model: estimate the density of the idiosyncratic shock to the running variable, and if it shows no periodicity of length $h-l$, the constant-bias assumption is likely violated.","Because the upper bound rests on an untestable monotonicity condition, an external validation study that observes control outcomes for lower-cutoff units just above their cutoff would provide a direct check on the headline bounds.","The pointwise-sharp-but-not-uniformly-sharp distinction suggests that developing uniform-sharp bounds under slightly stronger shape restrictions would be a useful next step."],"forward_implications":["When the running variable is non-manipulable and the groups are comparable in unobservables, the constant-bias extrapolation is justified.","When the running variable is partially manipulable, the constant-bias extrapolation can be substantially biased, and the proposed bounds remain valid without it.","The bounds are pointwise sharp, so no tighter interval can be obtained at a single point under the maintained assumptions.","Adding intermediate cutoff groups narrows the lower bound, so richer multi-cutoff designs give more informative identified sets.","The same bounding logic extends to one-sided fuzzy designs, covering settings in which eligible individuals may not take up the treatment."],"supporting_citations":[{"why":"Defines the constant-bias assumption and the extrapolation formula this paper scrutinizes and complements.","marker":"Cattaneo et al. (2021)"},{"why":"Provides the learning-in-games model of natural experiments on which Section 2's decision problem builds.","marker":"Fudenberg and Levine (2022)"},{"why":"Defines partial manipulability of the running variable and the density test used to motivate the failure case.","marker":"McCrary (2008)"},{"why":"Establishes the continuity-based identification of RD treatment effects at the cutoff, the baseline for extrapolation.","marker":"Hahn et al. (2001)"},{"why":"Supplies the empirical monotonicity of RD regression functions that motivates Assumption 3.1.","marker":"Babii and Kumar (2023)"},{"why":"Founds the monotone-response partial identification approach that the bounds extend.","marker":"Manski (1997)"},{"why":"Provides the robust bias-corrected local polynomial inference used for pointwise and uniform confidence bands.","marker":"Calonico et al. (2014, 2018)"},{"why":"The source of the ACCES data and design used in the second empirical illustration.","marker":"Melguizo et al. (2016)"},{"why":"The source of the SPP data and design used in the first empirical illustration.","marker":"Londoño-Vélez et al. (2020)"}],"fun_headline_variants":["Sharp bounds for RD extrapolation when cutoffs are gamed","Test scores break RD extrapolation; use bounds","Scores manipulable? RD extrapolation fails; bounds fix it","Bounding RD extrapolation bias from manipulable running variables"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the lower-cutoff group's untreated outcome function $\\mu_{0,l}(x)$ is monotone over the extrapolation interval, a restriction that is never observed and cannot be tested with regression-discontinuity data; if it fails, the upper bound collapses.","fun_headline_variants_meta":{"raw":{"variants":["Sharp bounds for RD extrapolation when cutoffs are gamed","Test scores break RD extrapolation; use bounds","Scores manipulable? RD extrapolation fails; bounds fix it","Bounding RD extrapolation bias from manipulable running variables"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00145,"raw_usage":{"total_tokens":5791,"prompt_tokens":850,"completion_tokens":4941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":4882}},"tokens_in":466,"tokens_out":4941,"duration_ms":33051,"temperature":1.0,"reasoning_tokens":4882,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:36:16.166994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the control-outcome function $\\mu_{0,l}(x)$ for $x \\in (l,h)$ using auxiliary data or a design that removes treatment for the lower-cutoff group; a non-monotone estimate invalidates the upper bound and the sharpness claim. For the decision-model failure result, estimate the density of score shocks across the two cutoff groups: an absence of $h-l$ periodicity indicates, by Proposition 2, that optimal effort responds to the cutoff and the constant-bias assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Founds the monotone-response partial identification approach that the bounds extend."},{"cited_title":"D., Keele, L., Titiunik, R., and Vazquez-Bare, G","cited_arxiv_id":null,"evidence_quote":"Defines the constant-bias assumption and the extrapolation formula this paper scrutinizes and complements."},{"cited_title":"and Levine, D","cited_arxiv_id":null,"evidence_quote":"Provides the learning-in-games model of natural experiments on which Section 2's decision problem builds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines partial manipulability of the running variable and the density test used to motivate the failure case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the continuity-based identification of RD treatment effects at the cutoff, the baseline for extrapolation."},{"cited_title":"and Kumar, R","cited_arxiv_id":null,"evidence_quote":"Supplies the empirical monotonicity of RD regression functions that motivates Assumption 3.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The source of the ACCES data and design used in the second empirical illustration."}],"review_version":1}