REVIEW 4 major objections 4 minor 109 references
Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A binary randomized experiment can reveal how many subjects defied the intervention.
desk verdict A sound design-based likelihood extension that deserves review; the main unresolved issue is an unproven contradiction with Copas and a few empirical claims that need proof or code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the design-based likelihood: for each candidate four-type count vector $\theta$, sum over all possible counts $i$ of always takers randomized into intervention of the product $\binom{\theta_{11}}{i}\binom{\theta_{10}}{x_{I1}-i}\binom{\theta_{01}}{\theta_{11}+\theta_{01}-x_{C1}-i}\binom{\theta_{00}}{m+x_{C1}+i-\theta_{11}-\theta_{01}-x_{I1}}$, normalized by the number of random assignments. This sum counts how many randomizations would produce the observed data under a given $\theta$; the paper calls that count entropy and divides by the total number of assignments to get a likelihood. It is the randomization design, completely randomized or Bernoulli, that makes the likelihood vary within the Fréchet set, because the same data can be reached in more ways when the people who reproduce the data belong to fewer types and are balanced across arms within each type.
What would settle it
Take a completely randomized experiment with six subjects, three per arm, and data showing two takeups in intervention and one in control; the paper's formula gives the likelihood of the joint distribution with four compliers and two defiers as $12/20$, higher than the $8/20$ for zero defiers. If an exact enumeration of all assignments produced equal likelihoods across the three distributions in this Fréchet set, the central claim would be falsified; equivalently, any finite-sample data realization for which the design-based likelihood is constant across the Fréchet set rather than U-shaped would contradict the claim.
Extended reading notes
Core claim
The central claim is that a completely randomized or Bernoulli-randomized experiment with a binary outcome identifies, through its design, a likelihood function over the sample's joint distribution of potential outcomes. Let $\theta=(\theta_{11},\theta_{10},\theta_{01},\theta_{00})$ count always takers, compliers, defiers, and never takers. The data $x=(x_{I1},x_{I0},x_{C1},x_{C0})$ are counts of takeup and no-takeup in each arm. The design-based likelihood is proportional to a sum of products of binomial coefficients over the unknown number $i$ of always takers assigned to intervention; this sum differs across $\theta$ values even when the marginals $\theta_{1\bullet}$ and $\theta_{\bullet1}$ are fixed. Consequently the likelihood varies with the number of defiers $\theta_{01}$ inside the Fréchet bounds, and its maximizer selects one four-type count vector. In every empirical Fréchet set the authors examined, the likelihood is U-shaped and peaks at the lower or upper bound on defiers; in the two published experiments, the rule reports zero defiers in one and 21 defiers, or 18% of the sample, in the other.
Load-bearing premise
The framework stands or falls on treating the sample as fixed: the potential outcomes of the $n$ subjects are fixed, the only randomness is the assignment mechanism, and no subject's outcome depends on another subject's assignment; if a researcher instead wants a statement about a population, the design-based likelihood is not directly about that target.
Editorial extensions
If this is right
- Within the estimated Fréchet bounds, the MLE counts defiers: when the estimated average effect is positive, the MLE includes defiers exactly when control takeup is below half and intervention takeup is above half, unless takeup is zero in control or full in intervention.
- The 95% smallest credible sets for defiers in both applications include zero and the estimated upper Fréchet bound, so the evidence is weak but not absent; in the organ-donation experiment, variation within the estimated bounds allows the rule to exclude the middle counts of 8 and 9 defiers.
- Under a uniform prior and a zero-one utility for correct guesses, the maximum likelihood rule is Bayes optimal and therefore admissible, and its Bayes expected utility exceeds both a rule that is uniform over each Fréchet set and a rule that imposes monotonicity, with the gain increasing in sample size.
- The rule gives an evidence-based route to monotonicity: when the MLE has zero defiers, as in the smoking-quit payment experiment, monotonicity appears as a data-supported simplification, while in the organ-donation experiment the MLE suggests the weaker assumption of no never takers.
- The counts of all four types are recovered jointly, and the MLE preserves the estimated average effect, so the four-type estimate is consistent with the usual summary statistic while adding the full distribution of effects.
Reading between the lines
- If the same design-based likelihood is derived for matched-pair, stratified, or permuted-block randomizations, the entropy logic could yield exact finite-sample distributions for test statistics and confidence intervals without simulation, a direction the paper lists as open.
- The rule's ability to label specific people as compliers or defiers under an assumption like "no never takers" suggests a practical targeting strategy: measure covariates of the labeled subjects and direct future interventions toward compliers and away from defiers.
- The paper's survey of published randomized trials implies a reporting standard: if journals required exact randomization procedure and arm-specific counts, design-based likelihoods could be recomputed for a large stock of existing experiments.
- Because the maximizer tends to sit at a Fréchet bound, applied work may want to report both the bound estimates and the MLE; the difference between them encodes how much the exact design, rather than the average effect, contributes to what can be said about heterogeneity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a design-based likelihood for the joint distribution of principal strata in a fixed experimental sample, using only a binary intervention, a binary outcome, and the randomization design. The likelihood is used to construct a maximum-likelihood decision rule for the counts of always takers, compliers, defiers, and never takers, with special attention to whether the estimate includes defiers. The authors show that this likelihood varies with the number of defiers within the Frechet set determined by the estimated marginals, provide Bayes optimality under a uniform prior and 0-1 utility, and apply the rule to two published experiments. The paper also contributes a visualization of the MLE over all data configurations for sample sizes 50 and 200 and an R package, dbmle.
Significance. If the empirical and computational claims are fully supported, this is a valuable demonstration that exact randomization-based likelihoods can contain finite-sample information about the composition of principal strata beyond the average effect. Appendix B carefully derives the likelihood and correctly recovers the Copas (1973) expression; the Bayes optimality proof in Appendix D.2 is standard and clearly presented. The two applications are well chosen, and the authors' emphasis on weak evidence and wide credible sets is honest and useful. However, the paper's central claims about the location and shape of the global likelihood maximum rest on unsupported assertions that are in direct tension with a published claim by Copas. Until that tension is resolved, the main advertised mechanism is not fully established.
major comments (4)
- [2.4.1] The unresolved contradiction with Copas (1973) is load-bearing. The text says Copas claims the likelihood is always maximized at a distribution preserving the direct estimates of the marginal distributions and having either the maximal or minimal number of defiers, and then states: "In practice, we often find that the likelihood is maximized at a distribution that does not preserve these marginal estimates." No proof, counterexample, or exhaustive computational record is supplied for either side. This matters because the abstract and Section 2.2 advertise the mechanism as variation within the Frechet set determined by the estimated marginals, while Section 2.3 defines the estimator as the global maximizer over all theta. If the global maximizer can leave the estimated Frechet set, then Figure 3 and the summary rule ("the MLE includes defiers if...") are not consequences of the within-Frechet variation described in Section 2.2. The authors should either prove or disprove Copas's claim, or explicitly redefine the decision rule as maximizing within the estimated Frechet set and restate the associated claims accordingly.
- [2.2] The statement "In all empirical examples we have considered, the likelihood maintains a U-shape within the estimated Frechet set and all other Frechet sets, and is maximal at either the minimum or maximum number of defiers" is an unrestricted empirical generalization with no supporting theorem or reproducible evidence. The U-shape is used in the discussion of Figure 2 to justify excluding middle defier counts from the 95% smallest credible set within the estimated Frechet set, and it is connected to the "maximal at a bound" pattern that drives the applications. A finite-sample proof or a precise characterization of the data configurations for which the U-shape holds is needed; as written, this is an unsupported claim that would be false if even one Frechet set exhibits a different pattern.
- [3] Under the assumption of no never takers, the individual labeling sentence in the Johnson and Goldstein application reverses the principal-strata classification. Observed non-takeup in the intervention arm reveals Y_I = 0; with theta_00 = 0 the type must be a defier (theta_01), not a complier. Observed non-takeup in the control arm reveals Y_C = 0; with theta_00 = 0 the type must be a complier (theta_10), not a defier. The sentence "all people in intervention who do not take up must be compliers, and all people who do not take up in control must be defiers" is therefore wrong, and the Pearl "necessary/sufficient cause" statements built on it are also incorrect. Please correct the labels and the causal interpretation.
- [2.4.1] The claim "In all even-sized samples up to 200, the MLE includes defiers when takeup is below half in control and above half in intervention, unless takeup is zero in control or full in intervention" is presented as a computational fact without a supporting proof or fully reproducible exhaustive enumeration. This pattern is a central output of the paper and is used in Section 3 to predict the results of the two applications. If it is intended as a theorem, a proof is needed; if it is an empirical finding, the authors should specify the exact grid, the treatment of ties, and provide the code and outputs that verify every sample size up to 200.
minor comments (4)
- [2.2, Eq. (4)] The phrase "we round them if necessary" is not a well-defined algorithm. Since the Frechet bounds and the grid search are count-based, clarify how non-integer estimated marginals are converted to counts and whether the decision rule ever depends on the rounding choice.
- [1.1] The statement "the randomization guarantees that the count of each type is the same in each arm" is imprecise; randomization does not guarantee such equality, and in the stylized example the equality follows from monotonicity together with the estimated marginals. Rephrase to avoid a factual error.
- [3] The phrase "was is a 'sufficient' cause" contains a typo, and the word "compilers" appears where "compliers" is intended in the individual labeling discussion.
- [Figure 2 and Table A.1] Make the distinction between credible sets computed from the full posterior over all theta and the normalized likelihood restricted to the estimated Frechet set more prominent in the captions; the text explains this, but the figure and table labels could mislead readers into conflating the two objects.
Circularity Check
No significant circularity: the design-based likelihood is derived from the randomization design and principal-strata definitions, not from the defier count it estimates.
full rationale
The central derivation chain is self-contained. Appendix B derives L(θ|x) from the randomization design (Bernoulli or completely randomized) and the principal-strata definitions; no parameter is fitted to the defier count and then renamed as a prediction. The MLE is a direct maximization of this likelihood, and Appendix D proves Bayes optimality under the stated 0-1 loss and uniform prior, which is an explicit assumption rather than a hidden fit. The applications use external published experiments, and the paper repeatedly cautions that evidence is weak: the 95% smallest credible sets for defiers include zero and the estimated upper Fréchet bound. The unresolved disagreement with Copas (1973) about whether the global maximizer preserves the estimated marginals is a correctness or robustness issue, not circularity, because the paper's claimed variation within Fréchet sets is a stated property of the likelihood, not an input to it. Minor self-citations (Kowalski 2019a,b; Christy and Kowalski 2024a,b) appear only as prior-draft footnotes and as future-work pointers in Section 5; they do not supply the load-bearing likelihood derivation or the empirical applications. Therefore no circular step can be exhibited.
Assumptions & free parameters
assumptions (5)
- domain assumption Stable unit treatment value assumption (SUTVA): each subject's outcome depends only on their own potential outcomes and assignment, ruling out interference.
- domain assumption Design-based paradigm: potential outcomes and sample composition are fixed; all randomness comes from the random assignment.
- domain assumption Uniform prior over all compositions of n subjects into four principal strata.
- domain assumption Randomization mechanism is known and exactly as modeled (Bernoulli or completely randomized).
- standard math The sampling-based likelihood (Appendix C) is the correct contrast for the claim that the traditional likelihood is flat.
Cite this review
Pith. "Pith review of Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect." pith.science (2026). https://pith.science/paper/VM56LFWF
@misc{pith2026241216352,
author = {Pith},
title = {Pith review of: Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect},
year = {2026},
howpublished = {\url{https://pith.science/paper/VM56LFWF}},
note = {Machine review of arXiv:2412.16352}
}
read the original abstract
Using only a binary intervention and outcome and the design of the randomization within an experiment, we construct a design-based likelihood of the joint distribution of potential outcomes in the sample -- the numbers of always takers, compliers, defiers, and never takers. We develop a visualization to show that samples with defiers can sometimes generate the data in more ways than samples without, yielding a higher likelihood. This likelihood can vary within the Frechet bounds, even though the traditional likelihood does not. Evidence is weak, but it exists, as we illustrate with health applications and our dbmle package.
Figures
Reference graph
Works this paper leans on
-
[1]
Abadie, A. (2002). Bootstrap tests for distributional treatment effects in instrumental variable models. Journal of the American Statistical Association\/ 97\/ (457), 284--292
2002
-
[2]
Abadie, A. (2003). Semiparametric instrumental variable estimation of treatment response models. Journal of econometrics\/ 113\/ (2), 231--263
2003
-
[3]
Athey, G
Abadie, A., S. Athey, G. W. Imbens, and J. M. Wooldridge (2020). Sampling-based versus design-based uncertainty in regression analysis. Econometrica\/ 88\/ (1), 265--296
2020
-
[4]
Cawley, J
Alsan, M., J. Cawley, J. Doyle, Joseph J, and N. Skelley (2025, January). Mean reversion in randomized controlled trials: Implications for program targeting and heterogeneous treatment effects. Working Paper 33369, National Bureau of Economic Research
2025
-
[5]
Angrist, J. D., G. W. Imbens, and D. B. Rubin (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association\/ 91\/ (434), 444--455
1996
-
[6]
Athey, S. and G. W. Imbens (2017). The econometrics of randomized experiments. In Handbook of Economic Field Experiments , Volume 1, pp.\ 73--140. Elsevier
2017
-
[7]
Bai, Y. (2022). Optimality of matched-pair designs in randomized controlled trials. American Economic Review\/ 112\/ (12), 3911--3940
2022
- [8]
Show all 109 references
-
[9]
Balke, A. and J. Pearl (1997). Bounds on treatment effects from studies with imperfect compliance. Journal of the American Statistical Association\/ 92\/ (439), 1171--1176
1997
-
[10]
Barnard, G. A. (1947). Significance Tests for 2 x 2 Tables . Biometrika\/ 34\/ (1-2), 123--138
1947
-
[11]
Imai, and Z
Ben-Michael, E., K. Imai, and Z. Jiang (2024). Policy learning with asymmetric counterfactual utilities. Journal of the American Statistical Association\/ , 1--14
2024
-
[12]
R., J.-L
Bernard, G. R., J.-L. Vincent, P.-F. Laterre, S. P. LaRosa, J.-F. Dhainaut, A. Lopez-Rodriguez, J. S. Steingrub, G. E. Garber, J. D. Helterbrand, E. W. Ely, and C. J. Fisher (2001). Efficacy and safety of recombinant human activated protein c for severe sepsis. New England Jou...
2001
-
[13]
Bj \"o rklund, A. and R. Moffitt (1987). The estimation of wage gains and welfare gains in self-selection models. The Review of Economics and Statistics\/ , 42--49
1987
-
[14]
Boole, G. (1854). Of statistical conditions. In An Investigation of the Laws of Thought: On Which Are Founded the Mathematical Theories of Logic and Probabilities , Chapter 19, pp.\ 295–319. Walton and Maberly
-
[15]
Canner, P. L. (1970). Selecting one of two treatments when the responses are dichotomous. Journal of the American Statistical Association\/ 65\/ (329), 293--306
1970
-
[16]
Chan, D. C., M. Gentzkow, and C. Yu (2022). Selection with variation in diagnostic skill: Evidence from radiologists. The Quarterly Journal of Economics\/ 137\/ (2), 729--783
2022
-
[17]
Christy, N. and A. E. Kowalski (2024a). Counting defiers in health care with a design-based likelihood for the joint distribution of potential outcomes. arXiv preprint arXiv:2412.16352d\/
2024 arXiv
-
[18]
Christy, N. and A. E. Kowalski (2024b). Starting small: Prioritizing safety over efficacy in randomized experiments using the exact finite sample likelihood. arXiv preprint arxiv:2407.18206\/
2024 arXiv
-
[19]
Copas, J. B. (1973). Randomization models for the matched and unmatched 2 x 2 tables . Biometrika\/ 60\/ (3), 467--476
1973
-
[20]
Cox, D. R. (1958). Planning of Experiments . New York, NY: Wiley
1958
-
[21]
Cui, Y. and S. Han (2023). Policy learning with distributional welfare. arXiv preprint arXiv:2311.15878\/
2023 arXiv
-
[22]
Dawid, A. P. and M. Musio (2022). Effects of causes and causes of effects. Annual Review of Statistics and Its Application\/ 9\/ (1), 261--287
2022
-
[23]
Dehejia, R. H. (2005). Program evaluation as a decision problem. Journal of Econometrics\/ 125\/ (1-2), 141--173
2005
-
[24]
Ding, P. and L. W. Miratrix (2019). Model-free causal inference of binary experimental data. Scandinavian Journal of Statistics\/ 46\/ (1), 200--214
2019
-
[25]
Fan, Y. and S. S. Park (2010). Sharp bounds on the distribution of treatment effects and their statistical inference. Econometric Theory\/ 26\/ (3), 931--951
2010
-
[26]
Ferguson, T. S. (1967). Mathematical Statistics: A Decision Theoretic Approach . Academic Press
1967
-
[27]
Fern \'a ndez, A. A., J. L. Montiel Olea, C. Qiu, J. Stoye, and S. Tinda (2024). Robust bayes treatment choice with partial identification. arXiv preprint arXiv:2408.11621\/
2024
-
[28]
(2017, December)
Ferrie, C. (2017, December). Statistical physics for babies . Baby university. Naperville, IL: Sourcebooks
2017
-
[29]
Fisher, R. (1935). Design of Experiments\/ (1st ed.). Edinburgh: Oliver and Boyd
1935
-
[30]
Frangakis, C. E. and D. B. Rubin (2002). Principal stratification in causal inference. Biometrics\/ 58\/ (1), 21--29
2002
-
[31]
Fr\'echet, M. (1957). Les tableaux de corrélation et les programmes linéaires. Revue de l'Institut International de Statistique / Review of the International Statistical Institute\/ 25\/ (1/3), 23--40
1957
-
[32]
Freedman, D. A. and R. A. Purves (1969). Bayes' method for bookies. The Annals of Mathematical Statistics\/ 40\/ (4), 1177--1186
1969
-
[33]
Gelman, A. and G. Imbens (2013, November). Why ask why? F orward causal inference and reverse causal questions. Working Paper 19614, National Bureau of Economic Research
2013
-
[34]
Gelman, A. and K. O’Rourke (2017). Attitudes toward amalgamating evidence in statistics. https://sites.stat.columbia.edu/gelman/research/unpublished/amalgamating4.pdf
2017
-
[35]
Gneezy, U. and A. Rustichini (2000). A fine is a price. The journal of legal studies\/ 29\/ (1), 1--17
2000
-
[36]
Golan, A. (2002). Information and entropy econometrics—editor's view. Journal of econometrics\/ 107\/ (1-2), 1--15
2002
-
[37]
Greenland, S. and J. M. Robins (1986). Identifiability, exchangeability, and epidemiological confounding. International journal of epidemiology\/ 15\/ (3), 413--419
1986
-
[38]
Mehta, and N
Guggenberger, P., N. Mehta, and N. Pavlov (2024). Minimax regret treatment rules with finite samples when a quantile is the object of interest. Technical report, The Pennsylvania State University
2024
-
[39]
Heckman, J. J., J. Smith, and N. Clements (1997). Making the most out of programme evaluations and social experiments: Accounting for heterogeneity in programme impacts. The Review of Economic Studies\/ 64\/ (4), 487--535
1997
-
[40]
Heckman, J. J. and E. J. Vytlacil (1999). Local instrumental variables and latent variable models for identifying and bounding treatment effects. Proceedings of the National Academy of Sciences\/ 96\/ (8), 4730--4734
1999
-
[41]
Hirano, K. (2008). Decision theory in econometrics. The New Palgrave Dictionary of Economics, 2nd Edition. Eds. S. Durlauf and Le Blume. Palgrave Macmillan\/
2008
-
[42]
Hirano, K. and J. R. Porter (2009). Asymptotics for statistical treatment rules. Econometrica\/ 77\/ (5), 1683--1701
2009
-
[43]
Hirano, K. and J. R. Porter (2020). Asymptotic analysis of statistical decision rules in econometrics. In Handbook of econometrics , Volume 7, pp.\ 283--354. Elsevier
2020
-
[44]
u r Angewandte Mathematik der Universit \
Hoeffding, W. (1940). Scale-invariant correlation theory. Schriften des Mathematischen Instituts und des Instituts f \"u r Angewandte Mathematik der Universit \"a t Berlin\/ 5\/ (3), 181--233. Translated by Dana Quade in The Collected Works of Wassily Hoeffding, ed. Fisher, N....
1940
-
[45]
Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association\/ 81\/ (396), 945--960
1986
-
[46]
Horowitz, J. L. and C. F. Manski (2000). Nonparametric analysis of randomized experiments with missing covariate and outcome data. Journal of the American statistical Association\/ 95\/ (449), 77--84
2000
-
[47]
Huber, M. and G. Mellace (2012, May). Relaxing monotonicity in the identification of local average treatment effects . Economics Working Paper Series 1212, University of St. Gallen, School of Economics and Political Science
2012
-
[48]
Huber, M. and G. Mellace (2015). Testing instrument validity for late identification based on inequality moment constraints. Review of Economics and Statistics\/ 97\/ (2), 398--411
2015
-
[49]
Imbens, G. W. (2020). Potential outcome and directed acyclic graph approaches to causality: Relevance for empirical practice in economics. Journal of Economic Literature\/ 58\/ (4), 1129--1179
2020
-
[50]
Imbens, G. W. and J. D. Angrist (1994). Identification and estimation of local average treatment effects. Econometrica\/ 62\/ (2), 467--475
1994
-
[51]
Imbens, G. W. and C. F. Manski (2004). Confidence intervals for partially identified parameters. Econometrica\/ 72\/ (6), 1845--1857
2004
-
[52]
Imbens, G. W. and D. B. Rubin (1997). Estimating outcome distributions for compliers in instrumental variables models. The Review of Economic Studies\/ 64\/ (4), 555--574
1997
-
[53]
Jaynes, E. T. (1957a). Information theory and statistical mechanics. Physical review\/ 106\/ (4), 620
1957
-
[54]
Jaynes, E. T. (1957b). Information theory and statistical mechanics. ii. Physical review\/ 108\/ (2), 171
1957
-
[55]
Jaynes, E. T. (1968). Prior probabilities. IEEE Transactions on systems science and cybernetics\/ 4\/ (3), 227--241
1968
-
[56]
Johnson, E. J. and D. Goldstein (2003). Do defaults save lives?
2003
-
[57]
Katz, L. F., J. R. Kling, J. B. Liebman, et al. (2001). Moving to opportunity in boston: Early results of a randomized mobility experiment. The Quarterly Journal of Economics\/ 116\/ (2), 607--654
2001
-
[58]
Kempthorne, O. (1952). Design and Analysis of Experiments . New York: Wiley
1952
-
[59]
Kessler, J. B. and A. E. Roth (2025). Increasing organ donor registration as a means to increase transplantation: an experiment with actual organ donor registrations. American Economic Journal: Economic Policy\/ 17\/ (2), 60--83
2025
-
[60]
Kitagawa, T. (2015). A test for instrument validity. Econometrica\/ 83\/ (5), 2043--2063
2015
-
[61]
Kitagawa, T. and A. Tetenov (2018). Who should be treated? empirical welfare maximization methods for treatment choice. Econometrica\/ 86\/ (2), 591--616
2018
-
[62]
Kline, P. M. and C. R. Walters (2020, March). Reasonable doubt: Experimental detection of job-level employment discrimination. Working Paper 26861, National Bureau of Economic Research
2020
-
[63]
Kowalski, A. E. (2019a, March). Counting defiers. Working Paper 25671, National Bureau of Economic Research
-
[64]
Kowalski, A. E. (2019b, March). A model of a randomized experiment with an application to the prowess clinical trial. Working Paper 25670, National Bureau of Economic Research
-
[65]
Kowalski, A. E. (2023a). Behaviour within a clinical trial and implications for mammography guidelines. The Review of Economic Studies\/ 90\/ (1), 432--462
2023
-
[66]
Kowalski, A. E. (2023b). Reconciling seemingly contradictory results from the O regon health insurance experiment and the M assachusetts health reform. The Review of Economics and Statistics\/ 105\/ (3), 646--664
2023
-
[67]
Kuhn, H. W. (1953). Extensive games and the problem of information. In H. W. Kuhn and A. W. Tucker (Eds.), Contributions to the Theory of Games, Volume II , pp.\ 193--216. Princeton: Princeton University Press
1953
-
[68]
Lacetera, N. and M. Macis (2010). Do all material incentives for pro-social activities backfire? the response to cash and non-cash incentives for blood donations. Journal of Economic Psychology\/ 31\/ (4), 738--748
2010
-
[69]
Li, A. and J. Pearl (2019). Unit selection based on counterfactual logic. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
2019
-
[70]
Li, X. and P. Ding (2016). Exact confidence intervals for the average causal effect on a binary outcome. Statistics in Medicine\/ 35\/ (6), 957--960
2016
-
[71]
Machado, C., A. M. Shaikh, and E. J. Vytlacil (2019). Instrumental variables and the sign of the average treatment effect. Journal of Econometrics\/ 212 , 522--555
2019
-
[72]
Manski, C. F. (1997a). The mixing problem in programme evaluation. The Review of Economic Studies\/ 64\/ (4), 537--553
1997
-
[73]
Manski, C. F. (1997b). Monotone treatment response. Econometrica\/ 65\/ (6), 1311--1334
1997
-
[74]
Manski, C. F. (2004). Statistical treatment rules for heterogeneous populations. Econometrica\/ 72\/ (4), 1221--1246
2004
-
[75]
Manski, C. F. (2007). Minimax-regret treatment choice with missing outcome data. Journal of Econometrics\/ 139\/ (1), 105--115
2007
-
[76]
Manski, C. F. (2018). Reasonable patient care under uncertainty. Health Economics\/ 27\/ (10), 1397--1421
2018
-
[77]
Manski, C. F. (2019). Treatment choice with trial data: Statistical decision theory should supplant hypothesis testing. The American Statistician\/ 73\/ (sup1), 296--304
2019
-
[78]
Manski, C. F., G. D. Sandefur, S. McLanahan, and D. Powers (1992). Alternative estimates of the effect of family structure during adolescence on high school graduation. Journal of the American Statistical Association\/ 87\/ (417), 25--37
1992
-
[79]
Manski, C. F. and A. Tetenov (2007). Admissible treatment rules for a risk-averse planner with experimental data on an innovation. Journal of Statistical Planning and Inference\/ 137\/ (6), 1998--2010
2007
-
[80]
Manski, C. F. and A. Tetenov (2021). Statistical decision properties of imprecise trials assessing coronavirus disease 2019 (covid-19) drugs. Value in Health\/ 24\/ (5), 641--647
2021
-
[81]
Mellstr \"o m, C. and M. Johannesson (2008). Crowding out in blood donation: was titmuss right? Journal of the European Economic Association\/ 6\/ (4), 845--863
2008
-
[82]
Mourifi \'e , I. and Y. Wan (2017). Testing local average treatment effect assumptions. Review of Economics and Statistics\/ 99\/ (2), 305--313
2017
-
[83]
Mullahy, J. (2018). Individual results may vary: Inequality-probability bounds for some health-outcome treatment effects. Journal of Health Economics\/ 61 , 151 -- 162
2018
-
[84]
Neyman, J. (1923). On the application of probability theory to agricultural experiments. E ssay on principles. S ection 9. Roczniki Nauk Rolniczych\/ 10 , 1--51. Translated by D.M. Dabrowski and T.P. Speed in Statistical Science 5(4), pp. 465--472, 1990
1923
-
[85]
Pearl, J. (1999). Probabilities of causation: Three counterfactual interpretations and their identification. Synthese\/ 121\/ (1/2), 93--149
1999
-
[86]
Pearl, J. and D. Mackenzie (2018). The Book of Why: The New Science of Cause and Effect . Basic books
2018
-
[87]
Permutt, T. and J. R. Hebel (1989). Simultaneous-equation estimation in a clinical trial of the effect of smoking on birth weight. Biometrics\/ , 619--622
1989
-
[88]
Richardson, T. S. and J. M. Robins (2010). Analysis of the binary instrumental variable model. Heuristics, Probability and Causality: A Tribute to Judea Pearl\/ , 415--444
2010
-
[89]
Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology\/ 66\/ (5), 688--701
1974
-
[90]
Rubin, D. B. (1977). Assignment to treatment group on the basis of a covariate. Journal of Educational and Behavioral Statistics\/ 2\/ (1), 1--26
1977
-
[91]
Rubin, D. B. (1980). Randomization analysis of experimental data: The Fisher Randomization Test c omment. Journal of the American Statistical Association\/ 75\/ (371), 591--593
1980
-
[92]
Schlag, K. H. (2003). How to minimize maximum regret in repeated decision making. Unpublished Manuscript, European University Institute\/
2003
-
[93]
Schlag, K. H. (2007). Eleven - designing randomized experiments under minimax regret. Unpublished manuscript, European University Institute, Florence\/
2007
-
[94]
Schneider, F. H., P. Campos-Mercade, S. Meier, D. Pope, E. Wengstr \"o m, and A. N. Meier (2023). Financial incentives for vaccination do not have negative unintended consequences. Nature\/ 613\/ (7944), 526--533
2023
-
[95]
Semenova, V. (2024). Aggregated intersection bounds and aggregated minimax values. arXiv preprint arXiv:2303.00982\/
2024 arXiv
-
[96]
Stoye, J. (2007). Minimax regret treatment choice with incomplete data and many treatments. Econometric Theory\/ 23\/ (1), 190--199
2007
-
[97]
Stoye, J. (2009). Minimax regret treatment choice with finite samples. Journal of Econometrics\/ 151\/ (1), 70--81
2009
-
[98]
Stoye, J. (2012). Minimax regret treatment choice with covariates or with limited validity of experiments. Journal of Econometrics\/ 166\/ (1), 138--156
2012
-
[99]
Chernozhukov, and H
Tamer, E., V. Chernozhukov, and H. Hong (2004). Parameter set inference in a class of econometric models. In Econometric Society 2004 North American Winter Meetings , Number 382. Econometric Society
2004
-
[100]
Bauld, D
Tappin, D., L. Bauld, D. Purves, K. Boyd, L. Sinclair, S. MacAskill, J. McKell, B. Friel, A. McConnachie, L. De Caestecker, et al. (2015). Financial incentives for smoking cessation in pregnancy: randomised controlled trial. Bmj\/ 350
2015
-
[101]
Tchetgen Tchetgen, E. J. (2024). The nudge average treatment effect. arXiv preprint arxiv:2410.23590\/
2024 arXiv
-
[102]
Tetenov, A. (2012). Statistical treatment choice based on asymmetric minimax regret criteria. Journal of Econometrics\/ 166\/ (1), 157--165
2012
-
[103]
Tian, J. and J. Pearl (2000). Probabilities of causation: Bounds and identification. Annals of Mathematics and Artificial Intelligence\/ 28\/ (1-4), 287--313
2000
-
[104]
Wager, S. and S. Athey (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association\/ 113\/ (523), 1228--1242
2018
-
[105]
Wald, A. (1949). Statistical Decision Functions . The Annals of Mathematical Statistics\/ 20\/ (2), 165 -- 205
1949
-
[106]
Wang, J. L., T. Tran, and F. Abebe (2016). Maximum entropy and bayesian inference for the monty hall problem. Journal of Applied Mathematics and Physics\/ 4\/ (7), 1222--1230
2016
-
[107]
Welch, B. L. (1937). On the z-test in randomized blocks and latin squares. Biometrika\/ 29\/ (1/2), 21--52
1937
-
[108]
Young, A. (2019). Channeling Fisher: randomization tests and the statistical insignificance of seemingly significant experimental results. The Quarterly Journal of Economics\/ 134\/ (2), 557--598
2019
-
[109]
Zhang, J. L. and D. B. Rubin (2003). Estimation of causal effects via principal stratification when some outcomes are truncated by ``death". Journal of Educational and Behavioral Statistics\/ 28\/ (4), 353--368
2003
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.