Pith. sign in

REVIEW 4 major objections 4 minor 109 references

Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A binary randomized experiment can reveal how many subjects defied the intervention.

desk verdict A sound design-based likelihood extension that deserves review; the main unresolved issue is an unproven contradiction with Copas and a few empirical claims that need proof or code. read the letter →

arxiv 2412.16352 v6 pith:VM56LFWF submitted 2024-12-20 econ.EM

classification econ.EM
keywords design-basedinferencerandomizedexperimentspotentialoutcomesprincipalstratadefierscompliersFréchetboundsmaximumlikelihood
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that even with only a binary intervention and a binary outcome, the experiment's randomization design itself supplies information about the joint distribution of potential outcomes in the sample—specifically the counts of always takers, compliers, defiers, and never takers. The design-based likelihood, which averages over the unobserved allocation of these types into arms, varies with the number of defiers within the Fréchet bounds set by the estimated marginal takeup rates, whereas the traditional sampling-based likelihood is flat across every such set. Maximizing this likelihood yields an estimate of the full four-type distribution, not just an average effect. The point matters because a positive average effect can coexist with defiers, and the rule gives applied researchers a way to see whether the data lean toward monotonicity or toward a specific alternative such as "no never takers."

What carries the argument

The engine is the design-based likelihood: for each candidate four-type count vector $\theta$, sum over all possible counts $i$ of always takers randomized into intervention of the product $\binom{\theta_{11}}{i}\binom{\theta_{10}}{x_{I1}-i}\binom{\theta_{01}}{\theta_{11}+\theta_{01}-x_{C1}-i}\binom{\theta_{00}}{m+x_{C1}+i-\theta_{11}-\theta_{01}-x_{I1}}$, normalized by the number of random assignments. This sum counts how many randomizations would produce the observed data under a given $\theta$; the paper calls that count entropy and divides by the total number of assignments to get a likelihood. It is the randomization design, completely randomized or Bernoulli, that makes the likelihood vary within the Fréchet set, because the same data can be reached in more ways when the people who reproduce the data belong to fewer types and are balanced across arms within each type.

What would settle it

Take a completely randomized experiment with six subjects, three per arm, and data showing two takeups in intervention and one in control; the paper's formula gives the likelihood of the joint distribution with four compliers and two defiers as $12/20$, higher than the $8/20$ for zero defiers. If an exact enumeration of all assignments produced equal likelihoods across the three distributions in this Fréchet set, the central claim would be falsified; equivalently, any finite-sample data realization for which the design-based likelihood is constant across the Fréchet set rather than U-shaped would contradict the claim.

Watch

Extended reading notes

Core claim

The central claim is that a completely randomized or Bernoulli-randomized experiment with a binary outcome identifies, through its design, a likelihood function over the sample's joint distribution of potential outcomes. Let $\theta=(\theta_{11},\theta_{10},\theta_{01},\theta_{00})$ count always takers, compliers, defiers, and never takers. The data $x=(x_{I1},x_{I0},x_{C1},x_{C0})$ are counts of takeup and no-takeup in each arm. The design-based likelihood is proportional to a sum of products of binomial coefficients over the unknown number $i$ of always takers assigned to intervention; this sum differs across $\theta$ values even when the marginals $\theta_{1\bullet}$ and $\theta_{\bullet1}$ are fixed. Consequently the likelihood varies with the number of defiers $\theta_{01}$ inside the Fréchet bounds, and its maximizer selects one four-type count vector. In every empirical Fréchet set the authors examined, the likelihood is U-shaped and peaks at the lower or upper bound on defiers; in the two published experiments, the rule reports zero defiers in one and 21 defiers, or 18% of the sample, in the other.

Load-bearing premise

The framework stands or falls on treating the sample as fixed: the potential outcomes of the $n$ subjects are fixed, the only randomness is the assignment mechanism, and no subject's outcome depends on another subject's assignment; if a researcher instead wants a statement about a population, the design-based likelihood is not directly about that target.

Editorial extensions

If this is right

  • Within the estimated Fréchet bounds, the MLE counts defiers: when the estimated average effect is positive, the MLE includes defiers exactly when control takeup is below half and intervention takeup is above half, unless takeup is zero in control or full in intervention.
  • The 95% smallest credible sets for defiers in both applications include zero and the estimated upper Fréchet bound, so the evidence is weak but not absent; in the organ-donation experiment, variation within the estimated bounds allows the rule to exclude the middle counts of 8 and 9 defiers.
  • Under a uniform prior and a zero-one utility for correct guesses, the maximum likelihood rule is Bayes optimal and therefore admissible, and its Bayes expected utility exceeds both a rule that is uniform over each Fréchet set and a rule that imposes monotonicity, with the gain increasing in sample size.
  • The rule gives an evidence-based route to monotonicity: when the MLE has zero defiers, as in the smoking-quit payment experiment, monotonicity appears as a data-supported simplification, while in the organ-donation experiment the MLE suggests the weaker assumption of no never takers.
  • The counts of all four types are recovered jointly, and the MLE preserves the estimated average effect, so the four-type estimate is consistent with the usual summary statistic while adding the full distribution of effects.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the same design-based likelihood is derived for matched-pair, stratified, or permuted-block randomizations, the entropy logic could yield exact finite-sample distributions for test statistics and confidence intervals without simulation, a direction the paper lists as open.
  • The rule's ability to label specific people as compliers or defiers under an assumption like "no never takers" suggests a practical targeting strategy: measure covariates of the labeled subjects and direct future interventions toward compliers and away from defiers.
  • The paper's survey of published randomized trials implies a reporting standard: if journals required exact randomization procedure and arm-specific counts, design-based likelihoods could be recomputed for a large stock of existing experiments.
  • Because the maximizer tends to sit at a Fréchet bound, applied work may want to report both the bound estimates and the MLE; the difference between them encodes how much the exact design, rather than the average effect, contributes to what can be said about heterogeneity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper develops a design-based likelihood for the joint distribution of principal strata in a fixed experimental sample, using only a binary intervention, a binary outcome, and the randomization design. The likelihood is used to construct a maximum-likelihood decision rule for the counts of always takers, compliers, defiers, and never takers, with special attention to whether the estimate includes defiers. The authors show that this likelihood varies with the number of defiers within the Frechet set determined by the estimated marginals, provide Bayes optimality under a uniform prior and 0-1 utility, and apply the rule to two published experiments. The paper also contributes a visualization of the MLE over all data configurations for sample sizes 50 and 200 and an R package, dbmle.

Significance. If the empirical and computational claims are fully supported, this is a valuable demonstration that exact randomization-based likelihoods can contain finite-sample information about the composition of principal strata beyond the average effect. Appendix B carefully derives the likelihood and correctly recovers the Copas (1973) expression; the Bayes optimality proof in Appendix D.2 is standard and clearly presented. The two applications are well chosen, and the authors' emphasis on weak evidence and wide credible sets is honest and useful. However, the paper's central claims about the location and shape of the global likelihood maximum rest on unsupported assertions that are in direct tension with a published claim by Copas. Until that tension is resolved, the main advertised mechanism is not fully established.

major comments (4)
  1. [2.4.1] The unresolved contradiction with Copas (1973) is load-bearing. The text says Copas claims the likelihood is always maximized at a distribution preserving the direct estimates of the marginal distributions and having either the maximal or minimal number of defiers, and then states: "In practice, we often find that the likelihood is maximized at a distribution that does not preserve these marginal estimates." No proof, counterexample, or exhaustive computational record is supplied for either side. This matters because the abstract and Section 2.2 advertise the mechanism as variation within the Frechet set determined by the estimated marginals, while Section 2.3 defines the estimator as the global maximizer over all theta. If the global maximizer can leave the estimated Frechet set, then Figure 3 and the summary rule ("the MLE includes defiers if...") are not consequences of the within-Frechet variation described in Section 2.2. The authors should either prove or disprove Copas's claim, or explicitly redefine the decision rule as maximizing within the estimated Frechet set and restate the associated claims accordingly.
  2. [2.2] The statement "In all empirical examples we have considered, the likelihood maintains a U-shape within the estimated Frechet set and all other Frechet sets, and is maximal at either the minimum or maximum number of defiers" is an unrestricted empirical generalization with no supporting theorem or reproducible evidence. The U-shape is used in the discussion of Figure 2 to justify excluding middle defier counts from the 95% smallest credible set within the estimated Frechet set, and it is connected to the "maximal at a bound" pattern that drives the applications. A finite-sample proof or a precise characterization of the data configurations for which the U-shape holds is needed; as written, this is an unsupported claim that would be false if even one Frechet set exhibits a different pattern.
  3. [3] Under the assumption of no never takers, the individual labeling sentence in the Johnson and Goldstein application reverses the principal-strata classification. Observed non-takeup in the intervention arm reveals Y_I = 0; with theta_00 = 0 the type must be a defier (theta_01), not a complier. Observed non-takeup in the control arm reveals Y_C = 0; with theta_00 = 0 the type must be a complier (theta_10), not a defier. The sentence "all people in intervention who do not take up must be compliers, and all people who do not take up in control must be defiers" is therefore wrong, and the Pearl "necessary/sufficient cause" statements built on it are also incorrect. Please correct the labels and the causal interpretation.
  4. [2.4.1] The claim "In all even-sized samples up to 200, the MLE includes defiers when takeup is below half in control and above half in intervention, unless takeup is zero in control or full in intervention" is presented as a computational fact without a supporting proof or fully reproducible exhaustive enumeration. This pattern is a central output of the paper and is used in Section 3 to predict the results of the two applications. If it is intended as a theorem, a proof is needed; if it is an empirical finding, the authors should specify the exact grid, the treatment of ties, and provide the code and outputs that verify every sample size up to 200.
minor comments (4)
  1. [2.2, Eq. (4)] The phrase "we round them if necessary" is not a well-defined algorithm. Since the Frechet bounds and the grid search are count-based, clarify how non-integer estimated marginals are converted to counts and whether the decision rule ever depends on the rounding choice.
  2. [1.1] The statement "the randomization guarantees that the count of each type is the same in each arm" is imprecise; randomization does not guarantee such equality, and in the stylized example the equality follows from monotonicity together with the estimated marginals. Rephrase to avoid a factual error.
  3. [3] The phrase "was is a 'sufficient' cause" contains a typo, and the word "compilers" appears where "compliers" is intended in the individual labeling discussion.
  4. [Figure 2 and Table A.1] Make the distinction between credible sets computed from the full posterior over all theta and the normalized likelihood restricted to the estimated Frechet set more prominent in the captions; the text explains this, but the figure and table labels could mislead readers into conflating the two objects.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the design-based likelihood is derived from the randomization design and principal-strata definitions, not from the defier count it estimates.

full rationale

The central derivation chain is self-contained. Appendix B derives L(θ|x) from the randomization design (Bernoulli or completely randomized) and the principal-strata definitions; no parameter is fitted to the defier count and then renamed as a prediction. The MLE is a direct maximization of this likelihood, and Appendix D proves Bayes optimality under the stated 0-1 loss and uniform prior, which is an explicit assumption rather than a hidden fit. The applications use external published experiments, and the paper repeatedly cautions that evidence is weak: the 95% smallest credible sets for defiers include zero and the estimated upper Fréchet bound. The unresolved disagreement with Copas (1973) about whether the global maximizer preserves the estimated marginals is a correctness or robustness issue, not circularity, because the paper's claimed variation within Fréchet sets is a stated property of the likelihood, not an input to it. Minor self-citations (Kowalski 2019a,b; Christy and Kowalski 2024a,b) appear only as prior-draft footnotes and as future-work pointers in Section 5; they do not supply the load-bearing likelihood derivation or the empirical applications. Therefore no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities and fits no free parameters. Its central claim depends on the design-based assumption that the sample is fixed and only assignment is random, on SUTVA, on a known randomization mechanism, and on a uniform prior for the decision-theoretic and Bayesian analyses. These are explicit modeling choices rather than hidden fits.

assumptions (5)
  • domain assumption Stable unit treatment value assumption (SUTVA): each subject's outcome depends only on their own potential outcomes and assignment, ruling out interference.
    Invoked in Section 2.1 to justify the observed outcome equation (1).
  • domain assumption Design-based paradigm: potential outcomes and sample composition are fixed; all randomness comes from the random assignment.
    Adopted in Section 2.1, 'all randomness in the experimental data X comes from the random assignment of subjects.'
  • domain assumption Uniform prior over all compositions of n subjects into four principal strata.
    Used in Section 2.4.2 and Appendix D to establish Bayes optimality of the MLE and to construct credible sets in Appendix E; stated as a flat Dirichlet prior.
  • domain assumption Randomization mechanism is known and exactly as modeled (Bernoulli or completely randomized).
    Required to write the likelihood in Equations (8) and (9); the paper uses the reported designs of the two applications.
  • standard math The sampling-based likelihood (Appendix C) is the correct contrast for the claim that the traditional likelihood is flat.
    Background result used to motivate the contribution; the multinomial and binomial likelihoods are standard.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect." pith.science (2026). https://pith.science/paper/VM56LFWF

@misc{pith2026241216352,
  author       = {Pith},
  title        = {Pith review of: Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VM56LFWF}},
  note         = {Machine review of arXiv:2412.16352}
}
read the original abstract

Using only a binary intervention and outcome and the design of the randomization within an experiment, we construct a design-based likelihood of the joint distribution of potential outcomes in the sample -- the numbers of always takers, compliers, defiers, and never takers. We develop a visualization to show that samples with defiers can sometimes generate the data in more ways than samples without, yielding a higher likelihood. This likelihood can vary within the Frechet bounds, even though the traditional likelihood does not. Evidence is weak, but it exists, as we illustrate with health applications and our dbmle package.

Figures

Figures reproduced from arXiv: 2412.16352 by the authors.

Figure 1
Figure 1. In an Experiment with Six People, the Likelihood Varies Among Joint [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Likelihood Varies with the Number of Defiers within the Estimated Fr´echet Set [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Visualization of Our Proposed Statistical Decision Rule: How the MLE Varies [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Under Optimality Conditions, For Increasing Sample Sizes, Performance of [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

109 extracted references · 69 canonical work pages

  1. [1]

    Abadie, A. (2002). Bootstrap tests for distributional treatment effects in instrumental variable models. Journal of the American Statistical Association\/ 97\/ (457), 284--292

  2. [2]

    Abadie, A. (2003). Semiparametric instrumental variable estimation of treatment response models. Journal of econometrics\/ 113\/ (2), 231--263

  3. [3]

    Athey, G

    Abadie, A., S. Athey, G. W. Imbens, and J. M. Wooldridge (2020). Sampling-based versus design-based uncertainty in regression analysis. Econometrica\/ 88\/ (1), 265--296

  4. [4]

    Cawley, J

    Alsan, M., J. Cawley, J. Doyle, Joseph J, and N. Skelley (2025, January). Mean reversion in randomized controlled trials: Implications for program targeting and heterogeneous treatment effects. Working Paper 33369, National Bureau of Economic Research

  5. [5]

    Angrist, J. D., G. W. Imbens, and D. B. Rubin (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association\/ 91\/ (434), 444--455

  6. [6]

    Athey, S. and G. W. Imbens (2017). The econometrics of randomized experiments. In Handbook of Economic Field Experiments , Volume 1, pp.\ 73--140. Elsevier

  7. [7]

    Bai, Y. (2022). Optimality of matched-pair designs in randomized controlled trials. American Economic Review\/ 112\/ (12), 3911--3940

  8. [8]

    Huang, S

    Bai, Y., S. Huang, S. Moon, A. M. Shaikh, and E. J. Vytlacil (2024). On the identifying power of monotonicity for average treatment effects. arXiv preprint arxiv:2405.14104\/

Show all 109 references
  1. [9]

    Balke, A. and J. Pearl (1997). Bounds on treatment effects from studies with imperfect compliance. Journal of the American Statistical Association\/ 92\/ (439), 1171--1176

  2. [10]

    Barnard, G. A. (1947). Significance Tests for 2 x 2 Tables . Biometrika\/ 34\/ (1-2), 123--138

  3. [11]

    Imai, and Z

    Ben-Michael, E., K. Imai, and Z. Jiang (2024). Policy learning with asymmetric counterfactual utilities. Journal of the American Statistical Association\/ , 1--14

  4. [12]

    R., J.-L

    Bernard, G. R., J.-L. Vincent, P.-F. Laterre, S. P. LaRosa, J.-F. Dhainaut, A. Lopez-Rodriguez, J. S. Steingrub, G. E. Garber, J. D. Helterbrand, E. W. Ely, and C. J. Fisher (2001). Efficacy and safety of recombinant human activated protein c for severe sepsis. New England Jou...

  5. [13]

    Bj \"o rklund, A. and R. Moffitt (1987). The estimation of wage gains and welfare gains in self-selection models. The Review of Economics and Statistics\/ , 42--49

  6. [14]

    Boole, G. (1854). Of statistical conditions. In An Investigation of the Laws of Thought: On Which Are Founded the Mathematical Theories of Logic and Probabilities , Chapter 19, pp.\ 295–319. Walton and Maberly

  7. [15]

    Canner, P. L. (1970). Selecting one of two treatments when the responses are dichotomous. Journal of the American Statistical Association\/ 65\/ (329), 293--306

  8. [16]

    Chan, D. C., M. Gentzkow, and C. Yu (2022). Selection with variation in diagnostic skill: Evidence from radiologists. The Quarterly Journal of Economics\/ 137\/ (2), 729--783

  9. [17]

    Christy, N. and A. E. Kowalski (2024a). Counting defiers in health care with a design-based likelihood for the joint distribution of potential outcomes. arXiv preprint arXiv:2412.16352d\/

  10. [18]

    Christy, N. and A. E. Kowalski (2024b). Starting small: Prioritizing safety over efficacy in randomized experiments using the exact finite sample likelihood. arXiv preprint arxiv:2407.18206\/

  11. [19]

    Copas, J. B. (1973). Randomization models for the matched and unmatched 2 x 2 tables . Biometrika\/ 60\/ (3), 467--476

  12. [20]

    Cox, D. R. (1958). Planning of Experiments . New York, NY: Wiley

  13. [21]

    Cui, Y. and S. Han (2023). Policy learning with distributional welfare. arXiv preprint arXiv:2311.15878\/

  14. [22]

    Dawid, A. P. and M. Musio (2022). Effects of causes and causes of effects. Annual Review of Statistics and Its Application\/ 9\/ (1), 261--287

  15. [23]

    Dehejia, R. H. (2005). Program evaluation as a decision problem. Journal of Econometrics\/ 125\/ (1-2), 141--173

  16. [24]

    Ding, P. and L. W. Miratrix (2019). Model-free causal inference of binary experimental data. Scandinavian Journal of Statistics\/ 46\/ (1), 200--214

  17. [25]

    Fan, Y. and S. S. Park (2010). Sharp bounds on the distribution of treatment effects and their statistical inference. Econometric Theory\/ 26\/ (3), 931--951

  18. [26]

    Ferguson, T. S. (1967). Mathematical Statistics: A Decision Theoretic Approach . Academic Press

  19. [27]

    Fern \'a ndez, A. A., J. L. Montiel Olea, C. Qiu, J. Stoye, and S. Tinda (2024). Robust bayes treatment choice with partial identification. arXiv preprint arXiv:2408.11621\/

  20. [28]

    (2017, December)

    Ferrie, C. (2017, December). Statistical physics for babies . Baby university. Naperville, IL: Sourcebooks

  21. [29]

    Fisher, R. (1935). Design of Experiments\/ (1st ed.). Edinburgh: Oliver and Boyd

  22. [30]

    Frangakis, C. E. and D. B. Rubin (2002). Principal stratification in causal inference. Biometrics\/ 58\/ (1), 21--29

  23. [31]

    Fr\'echet, M. (1957). Les tableaux de corrélation et les programmes linéaires. Revue de l'Institut International de Statistique / Review of the International Statistical Institute\/ 25\/ (1/3), 23--40

  24. [32]

    Freedman, D. A. and R. A. Purves (1969). Bayes' method for bookies. The Annals of Mathematical Statistics\/ 40\/ (4), 1177--1186

  25. [33]

    Gelman, A. and G. Imbens (2013, November). Why ask why? F orward causal inference and reverse causal questions. Working Paper 19614, National Bureau of Economic Research

  26. [34]

    Gelman, A. and K. O’Rourke (2017). Attitudes toward amalgamating evidence in statistics. https://sites.stat.columbia.edu/gelman/research/unpublished/amalgamating4.pdf

  27. [35]

    Gneezy, U. and A. Rustichini (2000). A fine is a price. The journal of legal studies\/ 29\/ (1), 1--17

  28. [36]

    Golan, A. (2002). Information and entropy econometrics—editor's view. Journal of econometrics\/ 107\/ (1-2), 1--15

  29. [37]

    Greenland, S. and J. M. Robins (1986). Identifiability, exchangeability, and epidemiological confounding. International journal of epidemiology\/ 15\/ (3), 413--419

  30. [38]

    Mehta, and N

    Guggenberger, P., N. Mehta, and N. Pavlov (2024). Minimax regret treatment rules with finite samples when a quantile is the object of interest. Technical report, The Pennsylvania State University

  31. [39]

    Heckman, J. J., J. Smith, and N. Clements (1997). Making the most out of programme evaluations and social experiments: Accounting for heterogeneity in programme impacts. The Review of Economic Studies\/ 64\/ (4), 487--535

  32. [40]

    Heckman, J. J. and E. J. Vytlacil (1999). Local instrumental variables and latent variable models for identifying and bounding treatment effects. Proceedings of the National Academy of Sciences\/ 96\/ (8), 4730--4734

  33. [41]

    Hirano, K. (2008). Decision theory in econometrics. The New Palgrave Dictionary of Economics, 2nd Edition. Eds. S. Durlauf and Le Blume. Palgrave Macmillan\/

  34. [42]

    Hirano, K. and J. R. Porter (2009). Asymptotics for statistical treatment rules. Econometrica\/ 77\/ (5), 1683--1701

  35. [43]

    Hirano, K. and J. R. Porter (2020). Asymptotic analysis of statistical decision rules in econometrics. In Handbook of econometrics , Volume 7, pp.\ 283--354. Elsevier

  36. [44]

    u r Angewandte Mathematik der Universit \

    Hoeffding, W. (1940). Scale-invariant correlation theory. Schriften des Mathematischen Instituts und des Instituts f \"u r Angewandte Mathematik der Universit \"a t Berlin\/ 5\/ (3), 181--233. Translated by Dana Quade in The Collected Works of Wassily Hoeffding, ed. Fisher, N....

  37. [45]

    Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association\/ 81\/ (396), 945--960

  38. [46]

    Horowitz, J. L. and C. F. Manski (2000). Nonparametric analysis of randomized experiments with missing covariate and outcome data. Journal of the American statistical Association\/ 95\/ (449), 77--84

  39. [47]

    Huber, M. and G. Mellace (2012, May). Relaxing monotonicity in the identification of local average treatment effects . Economics Working Paper Series 1212, University of St. Gallen, School of Economics and Political Science

  40. [48]

    Huber, M. and G. Mellace (2015). Testing instrument validity for late identification based on inequality moment constraints. Review of Economics and Statistics\/ 97\/ (2), 398--411

  41. [49]

    Imbens, G. W. (2020). Potential outcome and directed acyclic graph approaches to causality: Relevance for empirical practice in economics. Journal of Economic Literature\/ 58\/ (4), 1129--1179

  42. [50]

    Imbens, G. W. and J. D. Angrist (1994). Identification and estimation of local average treatment effects. Econometrica\/ 62\/ (2), 467--475

  43. [51]

    Imbens, G. W. and C. F. Manski (2004). Confidence intervals for partially identified parameters. Econometrica\/ 72\/ (6), 1845--1857

  44. [52]

    Imbens, G. W. and D. B. Rubin (1997). Estimating outcome distributions for compliers in instrumental variables models. The Review of Economic Studies\/ 64\/ (4), 555--574

  45. [53]

    Jaynes, E. T. (1957a). Information theory and statistical mechanics. Physical review\/ 106\/ (4), 620

  46. [54]

    Jaynes, E. T. (1957b). Information theory and statistical mechanics. ii. Physical review\/ 108\/ (2), 171

  47. [55]

    Jaynes, E. T. (1968). Prior probabilities. IEEE Transactions on systems science and cybernetics\/ 4\/ (3), 227--241

  48. [56]

    Johnson, E. J. and D. Goldstein (2003). Do defaults save lives?

  49. [57]

    Katz, L. F., J. R. Kling, J. B. Liebman, et al. (2001). Moving to opportunity in boston: Early results of a randomized mobility experiment. The Quarterly Journal of Economics\/ 116\/ (2), 607--654

  50. [58]

    Kempthorne, O. (1952). Design and Analysis of Experiments . New York: Wiley

  51. [59]

    Kessler, J. B. and A. E. Roth (2025). Increasing organ donor registration as a means to increase transplantation: an experiment with actual organ donor registrations. American Economic Journal: Economic Policy\/ 17\/ (2), 60--83

  52. [60]

    Kitagawa, T. (2015). A test for instrument validity. Econometrica\/ 83\/ (5), 2043--2063

  53. [61]

    Kitagawa, T. and A. Tetenov (2018). Who should be treated? empirical welfare maximization methods for treatment choice. Econometrica\/ 86\/ (2), 591--616

  54. [62]

    Kline, P. M. and C. R. Walters (2020, March). Reasonable doubt: Experimental detection of job-level employment discrimination. Working Paper 26861, National Bureau of Economic Research

  55. [63]

    Kowalski, A. E. (2019a, March). Counting defiers. Working Paper 25671, National Bureau of Economic Research

  56. [64]

    Kowalski, A. E. (2019b, March). A model of a randomized experiment with an application to the prowess clinical trial. Working Paper 25670, National Bureau of Economic Research

  57. [65]

    Kowalski, A. E. (2023a). Behaviour within a clinical trial and implications for mammography guidelines. The Review of Economic Studies\/ 90\/ (1), 432--462

  58. [66]

    Kowalski, A. E. (2023b). Reconciling seemingly contradictory results from the O regon health insurance experiment and the M assachusetts health reform. The Review of Economics and Statistics\/ 105\/ (3), 646--664

  59. [67]

    Kuhn, H. W. (1953). Extensive games and the problem of information. In H. W. Kuhn and A. W. Tucker (Eds.), Contributions to the Theory of Games, Volume II , pp.\ 193--216. Princeton: Princeton University Press

  60. [68]

    Lacetera, N. and M. Macis (2010). Do all material incentives for pro-social activities backfire? the response to cash and non-cash incentives for blood donations. Journal of Economic Psychology\/ 31\/ (4), 738--748

  61. [69]

    Li, A. and J. Pearl (2019). Unit selection based on counterfactual logic. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence

  62. [70]

    Li, X. and P. Ding (2016). Exact confidence intervals for the average causal effect on a binary outcome. Statistics in Medicine\/ 35\/ (6), 957--960

  63. [71]

    Machado, C., A. M. Shaikh, and E. J. Vytlacil (2019). Instrumental variables and the sign of the average treatment effect. Journal of Econometrics\/ 212 , 522--555

  64. [72]

    Manski, C. F. (1997a). The mixing problem in programme evaluation. The Review of Economic Studies\/ 64\/ (4), 537--553

  65. [73]

    Manski, C. F. (1997b). Monotone treatment response. Econometrica\/ 65\/ (6), 1311--1334

  66. [74]

    Manski, C. F. (2004). Statistical treatment rules for heterogeneous populations. Econometrica\/ 72\/ (4), 1221--1246

  67. [75]

    Manski, C. F. (2007). Minimax-regret treatment choice with missing outcome data. Journal of Econometrics\/ 139\/ (1), 105--115

  68. [76]

    Manski, C. F. (2018). Reasonable patient care under uncertainty. Health Economics\/ 27\/ (10), 1397--1421

  69. [77]

    Manski, C. F. (2019). Treatment choice with trial data: Statistical decision theory should supplant hypothesis testing. The American Statistician\/ 73\/ (sup1), 296--304

  70. [78]

    Manski, C. F., G. D. Sandefur, S. McLanahan, and D. Powers (1992). Alternative estimates of the effect of family structure during adolescence on high school graduation. Journal of the American Statistical Association\/ 87\/ (417), 25--37

  71. [79]

    Manski, C. F. and A. Tetenov (2007). Admissible treatment rules for a risk-averse planner with experimental data on an innovation. Journal of Statistical Planning and Inference\/ 137\/ (6), 1998--2010

  72. [80]

    Manski, C. F. and A. Tetenov (2021). Statistical decision properties of imprecise trials assessing coronavirus disease 2019 (covid-19) drugs. Value in Health\/ 24\/ (5), 641--647

  73. [81]

    Mellstr \"o m, C. and M. Johannesson (2008). Crowding out in blood donation: was titmuss right? Journal of the European Economic Association\/ 6\/ (4), 845--863

  74. [82]

    Mourifi \'e , I. and Y. Wan (2017). Testing local average treatment effect assumptions. Review of Economics and Statistics\/ 99\/ (2), 305--313

  75. [83]

    Mullahy, J. (2018). Individual results may vary: Inequality-probability bounds for some health-outcome treatment effects. Journal of Health Economics\/ 61 , 151 -- 162

  76. [84]

    Neyman, J. (1923). On the application of probability theory to agricultural experiments. E ssay on principles. S ection 9. Roczniki Nauk Rolniczych\/ 10 , 1--51. Translated by D.M. Dabrowski and T.P. Speed in Statistical Science 5(4), pp. 465--472, 1990

  77. [85]

    Pearl, J. (1999). Probabilities of causation: Three counterfactual interpretations and their identification. Synthese\/ 121\/ (1/2), 93--149

  78. [86]

    Pearl, J. and D. Mackenzie (2018). The Book of Why: The New Science of Cause and Effect . Basic books

  79. [87]

    Permutt, T. and J. R. Hebel (1989). Simultaneous-equation estimation in a clinical trial of the effect of smoking on birth weight. Biometrics\/ , 619--622

  80. [88]

    Richardson, T. S. and J. M. Robins (2010). Analysis of the binary instrumental variable model. Heuristics, Probability and Causality: A Tribute to Judea Pearl\/ , 415--444

  81. [89]

    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology\/ 66\/ (5), 688--701

  82. [90]

    Rubin, D. B. (1977). Assignment to treatment group on the basis of a covariate. Journal of Educational and Behavioral Statistics\/ 2\/ (1), 1--26

  83. [91]

    Rubin, D. B. (1980). Randomization analysis of experimental data: The Fisher Randomization Test c omment. Journal of the American Statistical Association\/ 75\/ (371), 591--593

  84. [92]

    Schlag, K. H. (2003). How to minimize maximum regret in repeated decision making. Unpublished Manuscript, European University Institute\/

  85. [93]

    Schlag, K. H. (2007). Eleven - designing randomized experiments under minimax regret. Unpublished manuscript, European University Institute, Florence\/

  86. [94]

    Schneider, F. H., P. Campos-Mercade, S. Meier, D. Pope, E. Wengstr \"o m, and A. N. Meier (2023). Financial incentives for vaccination do not have negative unintended consequences. Nature\/ 613\/ (7944), 526--533

  87. [95]

    Semenova, V. (2024). Aggregated intersection bounds and aggregated minimax values. arXiv preprint arXiv:2303.00982\/

  88. [96]

    Stoye, J. (2007). Minimax regret treatment choice with incomplete data and many treatments. Econometric Theory\/ 23\/ (1), 190--199

  89. [97]

    Stoye, J. (2009). Minimax regret treatment choice with finite samples. Journal of Econometrics\/ 151\/ (1), 70--81

  90. [98]

    Stoye, J. (2012). Minimax regret treatment choice with covariates or with limited validity of experiments. Journal of Econometrics\/ 166\/ (1), 138--156

  91. [99]

    Chernozhukov, and H

    Tamer, E., V. Chernozhukov, and H. Hong (2004). Parameter set inference in a class of econometric models. In Econometric Society 2004 North American Winter Meetings , Number 382. Econometric Society

  92. [100]

    Bauld, D

    Tappin, D., L. Bauld, D. Purves, K. Boyd, L. Sinclair, S. MacAskill, J. McKell, B. Friel, A. McConnachie, L. De Caestecker, et al. (2015). Financial incentives for smoking cessation in pregnancy: randomised controlled trial. Bmj\/ 350

  93. [101]

    Tchetgen Tchetgen, E. J. (2024). The nudge average treatment effect. arXiv preprint arxiv:2410.23590\/

  94. [102]

    Tetenov, A. (2012). Statistical treatment choice based on asymmetric minimax regret criteria. Journal of Econometrics\/ 166\/ (1), 157--165

  95. [103]

    Tian, J. and J. Pearl (2000). Probabilities of causation: Bounds and identification. Annals of Mathematics and Artificial Intelligence\/ 28\/ (1-4), 287--313

  96. [104]

    Wager, S. and S. Athey (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association\/ 113\/ (523), 1228--1242

  97. [105]

    Wald, A. (1949). Statistical Decision Functions . The Annals of Mathematical Statistics\/ 20\/ (2), 165 -- 205

  98. [106]

    Wang, J. L., T. Tran, and F. Abebe (2016). Maximum entropy and bayesian inference for the monty hall problem. Journal of Applied Mathematics and Physics\/ 4\/ (7), 1222--1230

  99. [107]

    Welch, B. L. (1937). On the z-test in randomized blocks and latin squares. Biometrika\/ 29\/ (1/2), 21--52

  100. [108]

    Young, A. (2019). Channeling Fisher: randomization tests and the statistical insignificance of seemingly significant experimental results. The Quarterly Journal of Economics\/ 134\/ (2), 557--598

  101. [109]

    Zhang, J. L. and D. B. Rubin (2003). Estimation of causal effects via principal stratification when some outcomes are truncated by ``death". Journal of Educational and Behavioral Statistics\/ 28\/ (4), 353--368

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.