REVIEW 3 major objections 8 minor 128 references
Robust confidence intervals for weighted estimands
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
The paper constructs minimax-bias estimators and uniformly valid confidence intervals for weighted estimands by bounding differences via parameter heterogeneity and weight distance.
T0 review reviewed 2026-07-09 challenge →
load-bearing objection Solid paper on robust inference for weighted estimands. Core theory is correct and genuinely useful. One practical gap (class enlargement) is real but minor and well-handled. the 3 major comments →
Robust Inference for Weighted Estimands
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central object is the sharp bound on the difference between two weighted estimands: it equals the product of the heterogeneity in parameters (the GLS residual standard deviation of the parameter vector) and the distance between weights (the standard deviation of the difference in estimators). This Cauchy-Schwarz-based decomposition separates the unknown (heterogeneity) from the contested (weight choice), allowing each to be handled independently. The heterogeneity is bounded above using a quantile-unbiased upper confidence bound derived from the noncentral chi-squared distribution of the GLS residual sum of squares. The weight disagreement is controlled by taking the maximum distance to
What carries the argument
Weighted estimands tau_w(theta) = w'theta; GLS-based heterogeneity measure H(theta) = sqrt(theta'Q theta); weight distance ||lambda - w||_Sigma; noncentral chi-squared inversion for heterogeneity UCB; folded normal critical values for robust CI; minimax-bias weights w* minimizing max distance; bounded variance class, truncated simplex class, covariate balance class
Load-bearing premise
The asymptotic validity requires that the estimated class of alternative weights asymptotically contains the target alternatives. In practice this is handled by using a slightly enlarged class for estimation and then suppressing the enlargement in reported results, but if the enlargement is too small, coverage fails; if too large, intervals are unnecessarily wide.
What would settle it
If the containment condition fails (the estimated class of alternatives does not asymptotically contain the target weights), the robust CI can undercover. More fundamentally, if the heterogeneity UCB is badly calibrated (e.g., due to failure of asymptotic normality or covariance matrix inconsistency), the Bonferroni coverage guarantee breaks down.
If this is right
- Researchers can report a single robust confidence interval alongside their conventional one, and readers who disagree with the baseline weights can still draw valid inferences at a known confidence level, formalizing what is currently an informal robustness-check exercise.
- The bounded-variance class result shows that GLS (precision-weighted) estimators are not just variance-efficient but also minimax-bias-optimal when the only consensus is that alternative weights should yield estimators of bounded precision, giving GLS a double-optimality that provides a principled default when researchers face ambiguity over weight choice.
- The breakdown-value framework (the smallest perturbation at which the robust CI includes a threshold like zero) gives practitioners a scalar summary of robustness analogous to stability concepts in other areas, making weight-sensitivity directly comparable across studies.
- The Project STAR application demonstrates that even a well-known randomized experiment can have conclusions sensitive to small weight perturbations, suggesting that external validity concerns are quantifiable and sometimes binding even when internal validity is secure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops robust inference procedures for weighted estimands—weighted averages of group-level parameters that arise in event studies, multisite experiments, and regression settings. The core idea is that different readers may prefer different weighting schemes, and conventional CIs for a baseline estimand may undercover for alternative estimands under effect heterogeneity. The author establishes a sharp Cauchy-Schwarz bound on the difference between any two weighted estimands, decomposing it into a heterogeneity measure (the GLS residual standard deviation of the parameter vector) and a distance measure between weight vectors (the standard deviation of the difference in estimators). Using this decomposition, the paper constructs (i) a minimax-bias robust estimator that minimizes the maximum distance to a class of alternative weights, and (ii) a robust confidence interval that achieves uniform coverage over a class of alternative estimands by combining a heterogeneity UCB (via noncentral chi-squared inversion) with a Bonferroni adjustment. The framework accommodates several practically relevant classes of alternatives (bounded variance, truncated simplex, covariate balance) and their intersections. The finite-sample normal-model results (Section 4, Appendix C) are extended to uniform asymptotic validity under asymptotically normal estimates, consistent covariance estimation, and estimated weights (Section 5, Appendix D). Two empirical applications—an event study on学校
Significance. The paper addresses a well-motivated and practically important problem: robustness of inference to the choice of weights in weighted estimands, which is central to ongoing debates in event studies and multisite experiments. The methodological contributions are substantial and well-executed. The Cauchy-Schwarz bound (Proposition 1) is clean and sharp. The heterogeneity UCB (Proposition 2) leverages standard noncentral chi-squared inversion (Pfanzagl, 1994) in a novel application. The Bonferroni-type robust CI (Proposition 6) provides a transparent coverage guarantee of 1-(alpha+beta). The uniform asymptotic extension (Propositions U1-U7, Appendix D) is carefully developed with appropriate assumptions (U1-U5), including a thoughtful treatment of the sample-versus-population heterogeneity distinction (Remark 5) and a subsequence-based proof handling both bounded and diverging normalized heterogeneity regimes. The containment condition (25) and its slack-based sufficient condition (27) are transparently handled. The two empirical applications are well-chosen and illustrate contrasting outcomes (robustness for the Peru internet event study; sensitivity for Project STAR). The framework's
major comments (3)
- Section 5.7, Remark 8: The practical guidance for the containment condition (25) acknowledges that one should use an enlarged class (e.g., r = r_0 + delta for delta > 0) to ensure asymptotic coverage for the target class, but then states that in implementation and empirical applications, 'I will suppress this caveat and talk about coverage as if delta = 0.' While the author argues that delta = 0.0001 leaves displayed CIs unchanged, this is not a formal guarantee. The paper would benefit from either (a) a brief sensitivity check in the empirical applications showing that the reported breakdown values are indeed unchanged at the displayed precision for small delta, or (b) a more explicit caveat in the application sections (Sections 7.1-7.2) that the reported coverage statements are for the target class with the understanding that a small enlargement has been suppressed. As currently stated
- Section 5.6, Proposition U6 and Remark 5: The decision to target the sample heterogeneity Ĥ_n(theta_n) (using estimated Σ̂_n) rather than the population heterogeneity H_n(theta_n) is well-motivated—the problematic √n(Ĥ_n - H_n) term is avoided. However, the practical implication is that the asymptotic coverage guarantee in Proposition U7 is for Ĥ_n(theta_n), not H_n(theta_n). The paper should clarify in Section 6 (Practical Implementation) that the feasible procedure's coverage guarantee is for the sample-heterogeneity object, and briefly discuss whether this distinction matters in typical empirical settings where Σ̂_n is close to Σ_n.
- Section 7.2, Project STAR application: The heterogeneity UCB is reported as η̃ = 16.214, which is over twice the baseline t-statistic of 6.714. This suggests very large inferred heterogeneity. It would strengthen the analysis to briefly decompose this into its components—how much of the heterogeneity UCB is driven by genuine cross-site ATE variation versus sampling noise in the site-level estimates. A simple diagnostic (e.g., comparing η̃ to what would be expected under homogeneous effects) would help readers calibrate whether the sensitivity result is driven by real heterogeneity or by the procedure being conservative.
minor comments (8)
- Section 2.2, Example (Bounded Variance), Eq. (3): The condition r >= sigma_min / sigma_w is stated but sigma_min is defined only later in the same equation. Consider defining sigma_min before its first appearance in the inequality.
- Section 3.2: The notation F_χ²(x; η) for the noncentral chi-squared CDF is introduced but the noncentrality parameter is η² (since H(θ) = η implies the noncentrality is η²). This is clarified in the text but could be made more explicit to avoid confusion.
- Section 4.2: The critical value function cv_{1-α}(b) is defined as the (1-α)-quantile of the folded normal |N(b,1)|, but it is only later noted (in Section 6.1) that it can be computed via the noncentral chi-squared distribution with one degree of freedom. This computational detail would be helpful earlier.
- Table 1: The heterogeneity UCB values of 0.00 for math at event times ℓ=0 and ℓ=6 should be briefly explained—these correspond to cases where F_χ²(H²(θ̂); 0) ≤ β, so the UCB is set to zero by definition.
- Section 7.1: The SA CIs in Figure 1 use plug-in standard errors rather than LNK's bootstrap standard errors (as noted in footnote 24). While this is reasonable, a brief remark on whether the choice of standard error affects the robustness conclusions would be useful.
- Appendix C.3, Proof of Proposition 3: The invariance argument via the Hunt-Stein theorem is standard but dense. A brief remark connecting the maximal invariant θ̂'Qθ̂ to the noncentral chi-squared family's monotone likelihood ratio property would improve readability.
- References: The paper cites several forthcoming or preprint works (e.g., Adusumilli 2026, Andrews and Chen 2025, Chernozhukov et al. 2025, Lau 2026, Sarfati and Vilfort 2026). Ensure these are updated with final publication details when available.
- Section 6.1, Recipe 1: Step 2 references Eq. (31) for constructing ŵ* and B̂^β_min(Λ̂), but the equation number is not visible in the rendered text. Verify that equation numbering is correct in the final version.
Simulated Author's Rebuttal
The referee recommends minor revision and finds the paper's contributions substantial and well-executed. We address all three major comments below. For Comment 1 (Remark 8, delta suppression), we will add an explicit caveat in the application sections and verify sensitivity to small delta. For Comment 2 (sample vs population heterogeneity, Remark 5), we will clarify in Section 6 that the feasible procedure's coverage guarantee is for the sample-heterogeneity object and discuss the practical implications. For Comment 3 (Project STAR heterogeneity decomposition), we will add a diagnostic comparing the heterogeneity UCB to what would be expected under homogeneous effects.
read point-by-point responses
-
Referee: Section 5.7, Remark 8: The practical guidance for the containment condition (25) acknowledges that one should use an enlarged class (e.g., r = r_0 + delta for delta > 0) to ensure asymptotic coverage for the target class, but then states that in implementation and empirical applications, 'I will suppress this caveat and talk about coverage as if delta = 0.' While the author argues that delta = 0.0001 leaves displayed CIs unchanged, this is not a formal guarantee. The paper would benefit from either (a) a brief sensitivity check in the empirical applications showing that the reported breakdown values are indeed unchanged at the displayed precision for small delta, or (b) a more explicit caveat in the application sections (Sections 7.1-7.2) that the reported coverage statements are for the target class with the understanding that a small enlargement has been suppressed.
Authors: The referee is correct that the current treatment of the delta suppression in Remark 8 is informal. We will adopt both suggested remedies. First, we will add an explicit caveat in Sections 7.1 and 7.2 noting that the reported coverage statements are for the target class, with the understanding that a small enlargement (delta > 0) has been suppressed per the asymptotic theory in Section 5.7. Second, we will include a brief sensitivity check in each application showing that the reported breakdown values (r*_l for the event study, epsilon* and c-bar_d for Project STAR) are unchanged at the displayed precision for delta values such as 0.0001 and 0.001. In the event study application, the bounded variance simplex class with r = 1 is the relevant target, and the enlargement r = 1 + delta leaves the breakdown values unchanged because the maximum distance function (equation 10) is continuous in r and the displayed precision (two decimal places) is coarse relative to delta. For Project STAR, the truncated simplex and covariate balance classes are parameterized by epsilon and c-bar, and the same continuity argument applies. We agree that making this explicit is better than asking the reader to verify it. revision: yes
-
Referee: Section 5.6, Proposition U6 and Remark 5: The decision to target the sample heterogeneity H-hat_n(theta_n) (using estimated Sigma-hat_n) rather than the population heterogeneity H_n(theta_n) is well-motivated—the problematic sqrt(n)(H-hat_n - H_n) term is avoided. However, the practical implication is that the asymptotic coverage guarantee in Proposition U7 is for H-hat_n(theta_n), not H_n(theta_n). The paper should clarify in Section 6 (Practical Implementation) that the feasible procedure's coverage guarantee is for the sample-heterogeneity object, and briefly discuss whether this distinction matters in typical empirical settings where Sigma-hat_n is close to Sigma_n.
Authors: We agree that this distinction should be made explicit in Section 6. The current draft discusses the sample-versus-population heterogeneity distinction in Remark 5 (Section 5.6), but the practical implications are not carried forward to Section 6. We will add a paragraph in Section 6.1 clarifying that the feasible procedure's coverage guarantee (via Propositions U6 and U7) is for the sample-heterogeneity object H-hat_n(theta_n), which uses the estimated covariance matrix Sigma-hat_n as the GLS weighting matrix. We will then discuss the practical relevance: when Sigma-hat_n is a consistent estimator of Sigma_n (as maintained by Assumption U2), the distinction between H-hat_n and H_n is asymptotically negligible in the bounded-heterogeneity regime (Case 1 in Appendix D.6), because the term sqrt(n)(H-hat_n^2 - H_n^2)/(2*H_n) is O_p(sqrt(n) * ||Sigma-hat_n - Sigma_n|| / H_n) = o_p(1) under the maintained rate condition. In the diverging-heterogeneity regime (Case 2), the distinction is also asymptotically negligible because the normal approximation to the noncentral chi-squared distribution dominates. We will note that in finite samples, the distinction could matter if the covariance matrix estimator is noisy, but this is already partially addressed by the conservative nature of the Bonferroni adjustment. revision: yes
-
Referee: Section 7.2, Project STAR application: The heterogeneity UCB is reported as eta-tilde = 16.214, which is over twice the baseline t-statistic of 6.714. This suggests very large inferred heterogeneity. It would strengthen the analysis to briefly decompose this into its components—how much of the heterogeneity UCB is driven by genuine cross-site ATE variation versus sampling noise in the site-level estimates. A simple diagnostic (e.g., comparing eta-tilde to what would be expected under homogeneous effects) would help readers calibrate whether the sensitivity result is driven by real heterogeneity or by the procedure being conservative.
Authors: This is a constructive suggestion. We will add a diagnostic to Section 7.2 that helps calibrate the magnitude of eta-tilde = 16.214. Specifically, we will compute the expected value of the heterogeneity UCB under the null of homogeneous effects (H_n(theta_n) = 0), which corresponds to the (1-beta)-quantile of the central chi-squared distribution with K-1 = 77 degrees of freedom, scaled appropriately. Under homogeneity, the statistic n * H-hat_n^2(theta-hat_n) follows a central chi-squared distribution with 77 degrees of freedom, so the expected heterogeneity UCB under homogeneity can be computed as the square root of the 95th percentile of chi^2_77 divided by sqrt(n). This provides a natural benchmark: if eta-tilde substantially exceeds this benchmark, it suggests genuine cross-site heterogeneity rather than pure sampling noise. We will report this benchmark value and interpret it. Based on our preliminary calculations, the 95th percentile of chi^2_77 is approximately 98.5, yielding a benchmark of sqrt(98.5/n) = sqrt(98.5/3783) approx 0.161, which when scaled by sqrt(n) gives approximately sqrt(98.5) approx 9.92. Since eta-tilde = 16.214 exceeds this benchmark, the inferred heterogeneity is not purely an artifact of sampling noise under the homogeneous null. We will also note that the procedure is designed to be conservative (it is a valid UCB, not a point estimate), so part of the gap between eta-tilde and the point estimate of heterogeneity reflects the confidence level 1-beta rather than genuine heterogeneity alone. revision: yes
Circularity Check
No circularity found. The derivation chain is self-contained and uses standard statistical arguments throughout.
full rationale
The paper's central derivation chain proceeds through three load-bearing steps, each of which is a standard statistical argument with no reduction to inputs by construction. (1) Proposition 1 applies Cauchy-Schwarz to bound |(λ-w)'θ| ≤ H(θ)·||λ-w||_Σ, using the annihilator matrix A to project onto the space orthogonal to 1. This is a direct inequality, not a definition disguised as a result. (2) Proposition 2 constructs the heterogeneity UCB by observing that θ̂'Qθ̂ ~ χ²_{K-1}(H(θ)) and inverting the CDF. The optimality claim (Proposition 3) cites Pfanzagl (1994), an external reference, for the monotone likelihood ratio and maximin testing argument—not a self-citation. (3) Proposition 6 combines the bias UCB (Proposition 4, which follows from Propositions 1 and 2) with a Bonferroni correction: on the event E_θ (prob ≥ 1-β), noncoverage reduces to P(w'(θ-θ̂)/σ_w > z_{1-α}) = α, yielding total noncoverage ≤ α+β. This is a standard Bonferroni argument. The asymptotic extensions (Propositions U1-U7) use subsequence arguments, continuous mapping theorem, and Slutsky-type bounds—again standard. The one self-citation (Sarfati and Vilfort, 2026, in footnote 17) concerns a tangential property about variance-efficient estimators being bias-optimal under misspecification, and is not load-bearing for any of the paper's main results. No 'prediction' or 'first-principles result' reduces to a fitted input or a self-cited ansatz.
Axiom & Free-Parameter Ledger
free parameters (5)
- alpha =
0.05
- beta =
0.05
- r =
1.0 (default)
- epsilon =
varies
- c_bar =
varies
axioms (5)
- domain assumption hat_theta ~ N(theta, Sigma) with known Sigma (Assumption 1)
- domain assumption Lambda is nonempty, compact, and convex (Assumption 2)
- domain assumption Uniform asymptotic normality of estimates (Assumption U1)
- domain assumption Uniform sqrt(n)-consistency of covariance estimator (Assumption U2)
- domain assumption Slater condition for inequality constraints (Assumption U5iii)
Cite this review
Pith. "Pith review of Robust Inference for Weighted Estimands." pith.science (2026). https://pith.science/paper/B7YQ6HAO
@misc{pith2026260707524,
author = {Pith},
title = {Pith review of: Robust Inference for Weighted Estimands},
year = {2026},
howpublished = {\url{https://pith.science/paper/B7YQ6HAO}},
note = {Machine review of arXiv:2607.07524}
}
read the original abstract
Researchers often conduct inference on weighted estimands, defined as weighted averages of group-level effects. Example settings include event studies with cohort-level effects and experiments with site-level effects. Under heterogeneous effects, different weighting schemes yield estimands with distinct empirical and policy interpretations, leading to ambiguity and disagreement over the choice of weights. I establish bounds on differences between weighted estimands and confidence bounds on effect heterogeneity, which I use to construct estimators that minimize worst-case bias and confidence intervals that are uniformly valid over classes of weighted estimands. I apply these methods to an event study in Lakdawala, Nakasone, and Kho (2023), which studies the effects of school-based internet access on test scores. I find that results are robust to broad classes of weights. I then apply the methods to Tennessee's Project STAR experiment and find that results are sensitive to small departures from baseline weights.
Figures
Reference graph
Works this paper leans on
-
[1]
Revisiting event study designs, with an application to the estimation of the marginal propensity to consume , author=. 2017 , institution=
work page 2017
-
[2]
American economic review , volume=
Two-way fixed effects estimators with heterogeneous treatment effects , author=. American economic review , volume=. 2020 , publisher=
work page 2020
-
[3]
Journal of econometrics , volume=
Difference-in-differences with variation in treatment timing , author=. Journal of econometrics , volume=. 2021 , publisher=
work page 2021
-
[4]
Journal of econometrics , volume=
Estimating dynamic treatment effects in event studies with heterogeneous treatment effects , author=. Journal of econometrics , volume=. 2021 , publisher=
work page 2021
-
[5]
Journal of econometrics , volume=
Difference-in-differences with multiple time periods , author=. Journal of econometrics , volume=. 2021 , publisher=
work page 2021
-
[6]
American Economic Review: Insights , volume=
Pretest with caution: Event-study estimates after testing for parallel trends , author=. American Economic Review: Insights , volume=. 2022 , publisher=
work page 2022
-
[7]
Two-stage differences in differences
Two-stage differences in differences , author=. arXiv preprint arXiv:2207.05943 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[8]
Journal of econometrics , volume=
Design-based analysis in difference-in-differences settings with staggered adoption , author=. Journal of econometrics , volume=. 2022 , publisher=
work page 2022
-
[9]
Journal of Political Economy Microeconomics , volume=
Efficient estimation for staggered rollout designs , author=. Journal of Political Economy Microeconomics , volume=. 2023 , publisher=
work page 2023
-
[10]
Journal of Econometrics , volume=
What’s trending in difference-in-differences? A synthesis of the recent econometrics literature , author=. Journal of Econometrics , volume=. 2023 , publisher=
work page 2023
-
[11]
Review of Economic Studies , volume=
Revisiting event-study designs: robust and efficient estimation , author=. Review of Economic Studies , volume=. 2024 , publisher=
work page 2024
- [12]
-
[13]
Journal of Political Economy , volume=
Ambulance Taxis: The Impact of Regulation and Litigation on Health-Care Fraud , author=. Journal of Political Economy , volume=. 2025 , publisher=
work page 2025
-
[14]
American Economic Journal: Applied Economics , volume=
Down to the Wire: Leveraging Technology to Improve Electric Utility Cost Recovery , author=. American Economic Journal: Applied Economics , volume=
-
[15]
Efficient Difference-in-Differences and Event Study Estimators
Efficient Difference-in-Differences and Event Study Estimators , author=. arXiv preprint arXiv:2506.17729 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[16]
Harvesting Differences-in-Differences and Event-Study Evidence , author=. 2025 , institution=
work page 2025
-
[17]
American Political Science Review , volume=
Causal panel analysis under parallel trends: lessons from a large reanalysis study , author=. American Political Science Review , volume=. 2026 , publisher=
work page 2026
-
[18]
American Economic Journal: Economic Policy , volume=
Dynamic impacts of school-based internet access on student learning: Evidence from Peruvian public primary schools , author=. American Economic Journal: Economic Policy , volume=. 2023 , publisher=
work page 2023
-
[19]
American Economic Review , volume=
Undergraduate gender diversity and the direction of scientific research , author=. American Economic Review , volume=. 2025 , publisher=
work page 2025
-
[20]
The central role of the propensity score in observational studies for causal effects , author=. Biometrika , volume=. 1983 , publisher=
work page 1983
-
[21]
Journal of Business & Economic Statistics , volume=
On using linear regressions in welfare economics , author=. Journal of Business & Economic Statistics , volume=. 1996 , publisher=
work page 1996
-
[22]
Estimating the Labor Market Impact of Voluntary Military Service Using Social Security Data on Military Applicants , author=. Econometrica , volume=
-
[23]
Journal of econometrics , volume=
Semiparametric instrumental variable estimation of treatment response models , author=. Journal of econometrics , volume=. 2003 , publisher=
work page 2003
-
[24]
Moving the goalposts: Addressing limited overlap in the estimation of average treatment effects by changing the estimand , author=. 2006 , publisher=
work page 2006
-
[25]
Journal of econometrics , volume=
Nonparametric IV estimation of local average treatment effects with covariates , author=. Journal of econometrics , volume=. 2007 , publisher=
work page 2007
-
[26]
Tennessee’s student teacher achievement ratio (STAR) project , author=. Harvard Dataverse , volume=
-
[27]
Dealing with limited overlap in estimation of average treatment effects , author=. Biometrika , volume=. 2009 , publisher=
work page 2009
-
[28]
Extrapolate-ing: External validity and overidentification in the late framework , author=. 2010 , institution=
work page 2010
-
[29]
Journal of the American Statistical Association , volume=
Nonparametric causal effects based on incremental propensity score interventions , author=. Journal of the American Statistical Association , volume=. 2019 , publisher=
work page 2019
-
[30]
arXiv preprint arXiv:2105.08766 , year=
Trading-off Bias and Variance in Stratified Experiments and in Matching Studies, Under a Boundedness Condition on the Magnitude of the Treatment Effect , author=. arXiv preprint arXiv:2105.08766 , year=
-
[31]
arXiv preprint arXiv:2206.10717 , year=
Marginal interventional effects , author=. arXiv preprint arXiv:2206.10717 , year=
-
[32]
Review of Economics and Statistics , volume=
Long-term care hospitals: A case study in waste , author=. Review of Economics and Statistics , volume=. 2023 , publisher=
work page 2023
-
[33]
American Economic Review , volume=
Contamination bias in linear regressions , author=. American Economic Review , volume=. 2024 , publisher=
work page 2024
-
[34]
Estimating Treatment Effects Under Bounded Heterogeneity
Estimating Treatment Effects Under Bounded Heterogeneity , author=. arXiv preprint arXiv:2510.05454 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[35]
Imbens, Guido W. and Angrist, Joshua D. , title =. Econometrica , year =. doi:10.2307/2951620 , url =
-
[36]
The Review of Economic Studies , volume=
The interpretation of instrumental variables estimators in simultaneous equations models with an application to the demand for fish , author=. The Review of Economic Studies , volume=. 2000 , publisher=
work page 2000
-
[37]
Structural equations, treatment effects, and econometric policy evaluation 1 , author=. Econometrica , volume=. 2005 , publisher=
work page 2005
-
[38]
Estimation in an Instrumental Variables Model with Treatment Effect Heterogeneity , institution =
Koles. Estimation in an Instrumental Variables Model with Treatment Effect Heterogeneity , institution =. 2013 , month =
work page 2013
-
[39]
Journal of Political Economy , volume=
Beyond LATE with a discrete instrument , author=. Journal of Political Economy , volume=. 2017 , publisher=
work page 2017
-
[40]
Journal of Econometrics , volume=
2SLS with multiple treatments , author=. Journal of Econometrics , volume=. 2024 , publisher=
work page 2024
-
[41]
Journal of Econometrics , volume=
Instrumental variable estimation with first-stage heterogeneity , author=. Journal of Econometrics , volume=. 2024 , publisher=
work page 2024
-
[42]
When Should We (Not) Interpret Linear IV Estimands as LATE?
When should we (not) interpret linear iv estimands as late? , author=. arXiv preprint arXiv:2011.06695 , year=
work page internal anchor Pith review Pith/arXiv arXiv 2011
-
[43]
Journal of Causal Inference , volume=
Instruments with heterogeneous effects: Bias, monotonicity, and localness , author=. Journal of Causal Inference , volume=. 2020 , publisher=
work page 2020
-
[44]
Improving Inference from Simple Instruments through Compliance Estimation
Improving inference from simple instruments through compliance estimation , author=. arXiv preprint arXiv:2108.03726 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[45]
American Economic Review , volume=
The causal interpretation of two-stage least squares with multiple instrumental variables , author=. American Economic Review , volume=. 2021 , publisher=
work page 2021
- [46]
-
[47]
Review of Economics and Statistics , volume=
Interpreting OLS estimands when treatment effects are heterogeneous: Smaller groups get larger weights , author=. Review of Economics and Statistics , volume=. 2022 , publisher=
work page 2022
-
[48]
American Economic Review , volume=
Judging judge fixed effects , author=. American Economic Review , volume=. 2023 , publisher=
work page 2023
-
[49]
American Economic Review: Insights , volume=
Interpreting TSLS Estimators in Information Provision Experiments , author=. American Economic Review: Insights , volume=. 2025 , publisher=
work page 2025
-
[50]
The quarterly journal of economics , volume=
Experimental estimates of education production functions , author=. The quarterly journal of economics , volume=. 1999 , publisher=
work page 1999
-
[51]
Journal of statistical planning and inference , volume=
Improving predictive inference under covariate shift by weighting the log-likelihood function , author=. Journal of statistical planning and inference , volume=. 2000 , publisher=
work page 2000
-
[52]
Journal of econometrics , volume=
Predicting the efficacy of future training programs using past experiences at other locations , author=. Journal of econometrics , volume=. 2005 , publisher=
work page 2005
-
[53]
Brookings papers on education policy , pages=
What have researchers learned from Project STAR? , author=. Brookings papers on education policy , pages=. 2006 , publisher=
work page 2006
-
[54]
American journal of epidemiology , volume=
Generalizing evidence from randomized clinical trials to target populations: the ACTG 320 trial , author=. American journal of epidemiology , volume=. 2010 , publisher=
work page 2010
-
[55]
Journal of the Royal Statistical Society Series A: Statistics in Society , volume=
The use of propensity scores to assess the generalizability of results from randomized trials , author=. Journal of the Royal Statistical Society Series A: Statistics in Society , volume=. 2011 , publisher=
work page 2011
-
[56]
Journal of the Royal Statistical Society Series A: Statistics in Society , volume=
From sample average treatment effect to population average treatment effect on the treated: combining experimental with observational studies to estimate population treatment effects , author=. Journal of the Royal Statistical Society Series A: Statistics in Society , volume=. 2015 , publisher=
work page 2015
-
[57]
The Quarterly journal of economics , volume=
Site selection bias in program evaluation , author=. The Quarterly journal of economics , volume=. 2015 , publisher=
work page 2015
-
[58]
American Journal of Political Science , volume=
Does regression produce representative estimates of causal effects? , author=. American Journal of Political Science , volume=. 2016 , publisher=
work page 2016
- [59]
-
[60]
The Annals of Applied Statistics , pages=
Sensitivity analysis for an unobserved moderator in RCT-to-target-population generalization of treatment effects , author=. The Annals of Applied Statistics , pages=. 2017 , publisher=
work page 2017
-
[61]
Journal of the American Statistical Association , volume=
Balancing covariates via propensity score weighting , author=. Journal of the American Statistical Association , volume=. 2018 , publisher=
work page 2018
-
[62]
A simple approximation for evaluating external validity bias , author=. Economics Letters , volume=. 2019 , publisher=
work page 2019
-
[63]
Journal of Business & Economic Statistics , volume=
External validity in fuzzy regression discontinuity designs , author=. Journal of Business & Economic Statistics , volume=. 2020 , publisher=
work page 2020
-
[64]
Journal of Business & Economic Statistics , volume=
From local to global: External validity in a fertility natural experiment , author=. Journal of Business & Economic Statistics , volume=. 2021 , publisher=
work page 2021
-
[65]
A brief review of domain adaptation , author=. Advances in data science and information engineering: proceedings from ICDATA 2020 and IKE 2020 , pages=. 2021 , publisher=
work page 2020
-
[66]
IEEE transactions on pattern analysis and machine intelligence , volume=
Domain generalization: A survey , author=. IEEE transactions on pattern analysis and machine intelligence , volume=. 2022 , publisher=
work page 2022
-
[67]
American economic review , volume=
The economic costs of conflict: A case study of the Basque Country , author=. American economic review , volume=. 2003 , publisher=
work page 2003
-
[68]
Journal of the American statistical Association , volume=
Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program , author=. Journal of the American statistical Association , volume=. 2010 , publisher=
work page 2010
-
[69]
Journal of economic literature , volume=
Using synthetic controls: Feasibility, data requirements, and methodological aspects , author=. Journal of economic literature , volume=. 2021 , publisher=
work page 2021
-
[70]
Journal of the American Statistical Association , volume=
A penalized synthetic control estimator for disaggregated data , author=. Journal of the American Statistical Association , volume=. 2021 , publisher=
work page 2021
-
[71]
arXiv preprint arXiv:2511.05870 , year=
Synthetic Parallel Trends , author=. arXiv preprint arXiv:2511.05870 , year=
-
[72]
Robust Estimation of a Location Parameter , author=. Ann. Math. Statist. , volume=
-
[73]
Journal of mathematical economics , volume=
Maxmin expected utility with non-unique prior , author=. Journal of mathematical economics , volume=. 1989 , publisher=
work page 1989
-
[74]
Asymptotics for statistical treatment rules , author=. Econometrica , volume=. 2009 , publisher=
work page 2009
-
[75]
New perspectives on statistical decisions under ambiguity , author=. Annu. Rev. Econ. , volume=. 2012 , publisher=
work page 2012
-
[76]
Set coverage and robust policy , author=. Economics Letters , volume=. 2012 , publisher=
work page 2012
-
[77]
Statistical decision theory and Bayesian analysis , author=. 2013 , publisher=
work page 2013
-
[78]
Who should be treated? empirical welfare maximization methods for treatment choice , author=. Econometrica , volume=. 2018 , publisher=
work page 2018
-
[79]
A model of scientific communication , author=. Econometrica , volume=. 2021 , publisher=
work page 2021
-
[80]
Econometrics for decision making: Building foundations sketched by Haavelmo and Wald , author=. Econometrica , volume=. 2021 , publisher=
work page 2021
This paper was first reviewed by glm-5.2 on July 9, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.