{"id":"caa6d419-bd5c-419d-9cce-345935b97ae2","arxiv_id":"2412.16352","paper_version":6,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A design-based likelihood, derived only from the randomization mechanism, can provide weak evidence about the number of defiers in an experimental sample, yielding a point estimate with wide credible sets.","lead":"This paper derives a likelihood from the randomization design of a binary experiment and shows it can distinguish, weakly, between samples with and without defiers. The authors use it to build a maximum likelihood estimate of the numbers of always takers, compliers, defiers, and never takers in a fixed sample.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unresolved contradiction of Copas (Section 2.4.1) is the load-bearing issue: if the global MLE sometimes leaves the estimated Frechet set, the advertised 'MLE within Frechet bounds' claim is not what the rule actually does, yet no counterexample or proof is supplied.","rationale":"The reader's weakest assumption points to the design-based fixed-sample framing and SUTVA; I do not regard that as the most load-bearing risk because the paper explicitly targets the sample and SUTVA is standard. The more dangerous spot is the unproved contradiction of Copas: it is the one place where the paper asserts a mathematical fact about the MLE without derivation, and the authors' own language ('in practice', 'in all applications') signals that it is only empirical. The likelihood formula itself and the variation within a fixed Frechet set appear correct; the n = 6 example checks out, so the mathematical core is not obviously wrong. But the advertised decision rule is the global MLE, and if the global MLE does not preserve the estimated marginals, then the abstract's 'within the Frechet bounds determined by the estimated marginal distributions' is not a property of the object being estimated. The proposed exhaustive-grid test settles this cleanly. Since the reader already issued a CONDITIONAL verdict and this concern supports that verdict rather than reversing it, I select UNCHANGED.","tokens_in":29028,"tokens_out":12148,"duration_ms":108260,"concrete_test":"Independently evaluate Eq. (2) by exhaustive grid search over all theta = (theta_11, theta_10, theta_01, theta_00) summing to n, for every possible data table (x_I1, x_I0, x_C1, x_C0) in completely randomized designs with even n from 4 to 40 (m = n/2), and also for the two application data tables. For each realization, record whether every global maximizer satisfies theta_11 + theta_10 = n * x_I1 / m and theta_11 + theta_01 = n * x_C1 / (n - m), and whether theta_01 is at a Frechet bound implied by those estimated margins. If any counterexample exists, print the smallest (n, x) and the maximizer; if none exists through n = 40, the 'often' claim should be retracted or replaced by a theorem.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the behavior of the global MLE, not only on the fact that the likelihood varies. In Section 2.4.1 the authors assert: 'Copas (1973) claims that the likelihood is always maximized at a joint distribution ... preserves the direct estimates of the marginal distributions ... In practice, we often find that the likelihood is maximized at a distribution that does not preserve these marginal estimates.' No proof or counterexample is given. If Copas is right, the paper's contradiction is an error and the MLE always lies in the estimated Frechet set at a bound, which would support the headline. If Copas is wrong, then for many data cells in Figure 3 the global maximizer is outside the estimated Frechet set, so the MLE's defier count is not pinned down by the 'variation within Frechet bounds' mechanism advertised in the abstract and Section 2.2. The statements 'in all applications we have considered' and 'in all even-sized samples up to 200' are empirical, not mathematical, and cannot settle a claimed theorem. This is the weakest load-bearing step because both the decision rule's interpretation and its advertised connection to Frechet bounds turn on it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a design-based likelihood for the joint distribution of principal strata in a fixed experimental sample, using only a binary intervention, a binary outcome, and the randomization design. The likelihood is used to construct a maximum-likelihood decision rule for the counts of always takers, compliers, defiers, and never takers, with special attention to whether the estimate includes defiers. The authors show that this likelihood varies with the number of defiers within the Frechet set determined by the estimated marginals, provide Bayes optimality under a uniform prior and 0-1 utility, and apply the rule to two published experiments. The paper also contributes a visualization of the MLE over all data configurations for sample sizes 50 and 200 and an R package, dbmle.","tokens_in":29234,"tokens_out":15459,"duration_ms":135723,"significance":"If the empirical and computational claims are fully supported, this is a valuable demonstration that exact randomization-based likelihoods can contain finite-sample information about the composition of principal strata beyond the average effect. Appendix B carefully derives the likelihood and correctly recovers the Copas (1973) expression; the Bayes optimality proof in Appendix D.2 is standard and clearly presented. The two applications are well chosen, and the authors' emphasis on weak evidence and wide credible sets is honest and useful. However, the paper's central claims about the location and shape of the global likelihood maximum rest on unsupported assertions that are in direct tension with a published claim by Copas. Until that tension is resolved, the main advertised mechanism is not fully established.","major_comments":[{"comment":"The unresolved contradiction with Copas (1973) is load-bearing. The text says Copas claims the likelihood is always maximized at a distribution preserving the direct estimates of the marginal distributions and having either the maximal or minimal number of defiers, and then states: \"In practice, we often find that the likelihood is maximized at a distribution that does not preserve these marginal estimates.\" No proof, counterexample, or exhaustive computational record is supplied for either side. This matters because the abstract and Section 2.2 advertise the mechanism as variation within the Frechet set determined by the estimated marginals, while Section 2.3 defines the estimator as the global maximizer over all theta. If the global maximizer can leave the estimated Frechet set, then Figure 3 and the summary rule (\"the MLE includes defiers if...\") are not consequences of the within-Frechet variation described in Section 2.2. The authors should either prove or disprove Copas's claim, or explicitly redefine the decision rule as maximizing within the estimated Frechet set and restate the associated claims accordingly.","section":"2.4.1"},{"comment":"The statement \"In all empirical examples we have considered, the likelihood maintains a U-shape within the estimated Frechet set and all other Frechet sets, and is maximal at either the minimum or maximum number of defiers\" is an unrestricted empirical generalization with no supporting theorem or reproducible evidence. The U-shape is used in the discussion of Figure 2 to justify excluding middle defier counts from the 95% smallest credible set within the estimated Frechet set, and it is connected to the \"maximal at a bound\" pattern that drives the applications. A finite-sample proof or a precise characterization of the data configurations for which the U-shape holds is needed; as written, this is an unsupported claim that would be false if even one Frechet set exhibits a different pattern.","section":"2.2"},{"comment":"Under the assumption of no never takers, the individual labeling sentence in the Johnson and Goldstein application reverses the principal-strata classification. Observed non-takeup in the intervention arm reveals Y_I = 0; with theta_00 = 0 the type must be a defier (theta_01), not a complier. Observed non-takeup in the control arm reveals Y_C = 0; with theta_00 = 0 the type must be a complier (theta_10), not a defier. The sentence \"all people in intervention who do not take up must be compliers, and all people who do not take up in control must be defiers\" is therefore wrong, and the Pearl \"necessary/sufficient cause\" statements built on it are also incorrect. Please correct the labels and the causal interpretation.","section":"3"},{"comment":"The claim \"In all even-sized samples up to 200, the MLE includes defiers when takeup is below half in control and above half in intervention, unless takeup is zero in control or full in intervention\" is presented as a computational fact without a supporting proof or fully reproducible exhaustive enumeration. This pattern is a central output of the paper and is used in Section 3 to predict the results of the two applications. If it is intended as a theorem, a proof is needed; if it is an empirical finding, the authors should specify the exact grid, the treatment of ties, and provide the code and outputs that verify every sample size up to 200.","section":"2.4.1"}],"minor_comments":[{"comment":"The phrase \"we round them if necessary\" is not a well-defined algorithm. Since the Frechet bounds and the grid search are count-based, clarify how non-integer estimated marginals are converted to counts and whether the decision rule ever depends on the rounding choice.","section":"2.2, Eq. (4)"},{"comment":"The statement \"the randomization guarantees that the count of each type is the same in each arm\" is imprecise; randomization does not guarantee such equality, and in the stylized example the equality follows from monotonicity together with the estimated marginals. Rephrase to avoid a factual error.","section":"1.1"},{"comment":"The phrase \"was is a 'sufficient' cause\" contains a typo, and the word \"compilers\" appears where \"compliers\" is intended in the individual labeling discussion.","section":"3"},{"comment":"Make the distinction between credible sets computed from the full posterior over all theta and the normalized likelihood restricted to the estimated Frechet set more prominent in the captions; the text explains this, but the figure and table labels could mislead readers into conflating the two objects.","section":"Figure 2 and Table A.1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the main reason this is major rather than minor is the unresolved Copas contradiction in Section 2.4.1. If the authors can provide a counterexample or proof, or restrict the decision rule to the estimated Frechet set and adjust the claims accordingly, the paper could be suitable after revision. The reversed individual labeling in Section 3 is a clear factual error that must also be fixed. I see no issue with the core Appendix B derivation or the Bayes optimality argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper shows that a fixed-sample design-based analysis of a 2x2 experimental table carries weak but real information about the number of defiers. The core derivation is correct, and the central claim holds: the likelihood does vary with the joint distribution of potential outcomes, even within the Frechet set determined by the observed marginals. That is a legitimate extension of Copas's likelihood, which the paper credits properly. The visualization in Figure 1 is actually helpful, and the two applications are honest—the credible sets include zero and the upper bound, which is exactly the right way to present weak evidence.\n\nThe soft spots are real but not fatal. The biggest one is Section 2.4.1. The authors assert that Copas's claim about the maximizer is often false, but they supply no counterexample and no proof. This matters because they later say the likelihood is maximized at the estimated upper or lower Frechet bound. If Copas is right, the contradiction is an error. If the authors are right, they need to show it. The stress-test note is on target here: the paper's abstract only promises that the likelihood varies, which is true, but several passages in the text come close to claiming the MLE is inside the estimated Frechet set. That ambiguity needs to be resolved in revision.\n\nRelatedly, the blanket statements about all even-sized samples up to 200 are empirical claims from exhaustive search, not theorems. They are probably correct for the grids they ran, but the paper presents them as general facts. That is fine if the code is available, but the abstract mentions a dbmle package that I do not see in the arXiv files. Shipping the code would make those claims reproducible.\n\nThe Bayes optimality argument is standard but clean. The sampling-based versus design-based distinction is handled carefully. The paper does not overclaim: it repeatedly warns that the evidence about defiers is weak. That honesty is a real strength.\n\nAll told, the central argument holds up. The paper deserves a serious referee who can force the Copas question into the open. I would send it out rather than desk reject. It is a genuine contribution to design-based causal inference and principal stratification, and it will be more valuable once the proof or counterexample and the code are in place.","headline":"A sound design-based likelihood extension that deserves review; the main unresolved issue is an unproven contradiction with Copas and a few empirical claims that need proof or code.","tokens_in":29773,"tokens_out":3378,"would_cite":true,"duration_ms":31996,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A binary randomized experiment can reveal how many subjects defied the intervention.","keywords":["design-based inference","randomized experiments","potential outcomes","principal strata","defiers","compliers","Fréchet bounds","maximum likelihood"],"falsifier":"Take a completely randomized experiment with six subjects, three per arm, and data showing two takeups in intervention and one in control; the paper's formula gives the likelihood of the joint distribution with four compliers and two defiers as $12/20$, higher than the $8/20$ for zero defiers. If an exact enumeration of all assignments produced equal likelihoods across the three distributions in this Fréchet set, the central claim would be falsified; equivalently, any finite-sample data realization for which the design-based likelihood is constant across the Fréchet set rather than U-shaped would contradict the claim.","tokens_in":28797,"feed_emoji":"🧪","tokens_out":7979,"duration_ms":64407,"temperature":0.7,"pith_summary":"The paper claims that even with only a binary intervention and a binary outcome, the experiment's randomization design itself supplies information about the joint distribution of potential outcomes in the sample—specifically the counts of always takers, compliers, defiers, and never takers. The design-based likelihood, which averages over the unobserved allocation of these types into arms, varies with the number of defiers within the Fréchet bounds set by the estimated marginal takeup rates, whereas the traditional sampling-based likelihood is flat across every such set. Maximizing this likelihood yields an estimate of the full four-type distribution, not just an average effect. The point matters because a positive average effect can coexist with defiers, and the rule gives applied researchers a way to see whether the data lean toward monotonicity or toward a specific alternative such as \"no never takers.\"","feed_headline":"Randomization alone can estimate how many people defy treatment","feed_subtitle":"A design-based likelihood recovers counts of always takers, compliers, defiers, and never takers from a binary experiment.","key_machinery":"The engine is the design-based likelihood: for each candidate four-type count vector $\\theta$, sum over all possible counts $i$ of always takers randomized into intervention of the product $\\binom{\\theta_{11}}{i}\\binom{\\theta_{10}}{x_{I1}-i}\\binom{\\theta_{01}}{\\theta_{11}+\\theta_{01}-x_{C1}-i}\\binom{\\theta_{00}}{m+x_{C1}+i-\\theta_{11}-\\theta_{01}-x_{I1}}$, normalized by the number of random assignments. This sum counts how many randomizations would produce the observed data under a given $\\theta$; the paper calls that count entropy and divides by the total number of assignments to get a likelihood. It is the randomization design, completely randomized or Bernoulli, that makes the likelihood vary within the Fréchet set, because the same data can be reached in more ways when the people who reproduce the data belong to fewer types and are balanced across arms within each type.","core_discovery":"The central claim is that a completely randomized or Bernoulli-randomized experiment with a binary outcome identifies, through its design, a likelihood function over the sample's joint distribution of potential outcomes. Let $\\theta=(\\theta_{11},\\theta_{10},\\theta_{01},\\theta_{00})$ count always takers, compliers, defiers, and never takers. The data $x=(x_{I1},x_{I0},x_{C1},x_{C0})$ are counts of takeup and no-takeup in each arm. The design-based likelihood is proportional to a sum of products of binomial coefficients over the unknown number $i$ of always takers assigned to intervention; this sum differs across $\\theta$ values even when the marginals $\\theta_{1\\bullet}$ and $\\theta_{\\bullet1}$ are fixed. Consequently the likelihood varies with the number of defiers $\\theta_{01}$ inside the Fréchet bounds, and its maximizer selects one four-type count vector. In every empirical Fréchet set the authors examined, the likelihood is U-shaped and peaks at the lower or upper bound on defiers; in the two published experiments, the rule reports zero defiers in one and 21 defiers, or 18% of the sample, in the other.","pith_inferences":["If the same design-based likelihood is derived for matched-pair, stratified, or permuted-block randomizations, the entropy logic could yield exact finite-sample distributions for test statistics and confidence intervals without simulation, a direction the paper lists as open.","The rule's ability to label specific people as compliers or defiers under an assumption like \"no never takers\" suggests a practical targeting strategy: measure covariates of the labeled subjects and direct future interventions toward compliers and away from defiers.","The paper's survey of published randomized trials implies a reporting standard: if journals required exact randomization procedure and arm-specific counts, design-based likelihoods could be recomputed for a large stock of existing experiments.","Because the maximizer tends to sit at a Fréchet bound, applied work may want to report both the bound estimates and the MLE; the difference between them encodes how much the exact design, rather than the average effect, contributes to what can be said about heterogeneity."],"forward_implications":["Within the estimated Fréchet bounds, the MLE counts defiers: when the estimated average effect is positive, the MLE includes defiers exactly when control takeup is below half and intervention takeup is above half, unless takeup is zero in control or full in intervention.","The 95% smallest credible sets for defiers in both applications include zero and the estimated upper Fréchet bound, so the evidence is weak but not absent; in the organ-donation experiment, variation within the estimated bounds allows the rule to exclude the middle counts of 8 and 9 defiers.","Under a uniform prior and a zero-one utility for correct guesses, the maximum likelihood rule is Bayes optimal and therefore admissible, and its Bayes expected utility exceeds both a rule that is uniform over each Fréchet set and a rule that imposes monotonicity, with the gain increasing in sample size.","The rule gives an evidence-based route to monotonicity: when the MLE has zero defiers, as in the smoking-quit payment experiment, monotonicity appears as a data-supported simplification, while in the organ-donation experiment the MLE suggests the weaker assumption of no never takers.","The counts of all four types are recovered jointly, and the MLE preserves the estimated average effect, so the four-type estimate is consistent with the usual summary statistic while adding the full distribution of effects."],"supporting_citations":[{"why":"Supplies the design-based likelihood for the completely randomized 2x2 table; equation (2) of the paper is proportional to Copas's expression.","marker":"Copas (1973)"},{"why":"Establishes the principal-strata classification of always takers, compliers, defiers, and never takers that the paper counts.","marker":"Angrist, Imbens, and Rubin (1996)"},{"why":"Defines the monotonicity assumption that the paper's rule can support or reject and supplies the baseline for the monotonicity decision rule.","marker":"Imbens and Angrist (1994)"},{"why":"Introduces potential outcomes and randomization-based inference, the foundation of the design-based model.","marker":"Neyman (1923)"},{"why":"Provides the potential-outcome causal model and assignment-independence condition the design-based likelihood uses.","marker":"Rubin (1974, 1977)"},{"why":"Derive the bounds on joint distributions given fixed margins that define the Fréchet set within which defier counts vary.","marker":"Boole (1854), Hoeffding (1940), and Fréchet (1957)"},{"why":"Supplies the maximum-entropy principle used to interpret the likelihood maximizer as the least informative distribution consistent with the data.","marker":"Jaynes (1957a,b)"},{"why":"Provides one of the two empirical applications: a completely randomized payment intervention for pregnant smokers with a positive average effect.","marker":"Tappin et al. (2015)"},{"why":"Provides the other empirical application: a Bernoulli-randomized organ-donation default experiment whose MLE includes 21 defiers.","marker":"Johnson and Goldstein (2003)"}],"fun_headline_variants":["Defiers detectable from randomization alone","Count defiers from design, not just averages","Beyond averages: the design of your experiment counts defiers","Randomization design reveals defier counts, not just averages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework stands or falls on treating the sample as fixed: the potential outcomes of the $n$ subjects are fixed, the only randomness is the assignment mechanism, and no subject's outcome depends on another subject's assignment; if a researcher instead wants a statement about a population, the design-based likelihood is not directly about that target.","fun_headline_variants_meta":{"raw":{"variants":["Defiers detectable from randomization alone","Count defiers from design, not just averages","Beyond averages: the design of your experiment counts defiers","Randomization design reveals defier counts, not just averages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1435,"prompt_tokens":903,"completion_tokens":532,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":472}},"tokens_in":519,"tokens_out":532,"duration_ms":5422,"temperature":1.0,"reasoning_tokens":472,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:39:44.972074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a completely randomized experiment with six subjects, three per arm, and data showing two takeups in intervention and one in control; the paper's formula gives the likelihood of the joint distribution with four compliers and two defiers as $12/20$, higher than the $8/20$ for zero defiers. If an exact enumeration of all assignments produced equal likelihoods across the three distributions in this Fréchet set, the central claim would be falsified; equivalently, any finite-sample data realization for which the design-based likelihood is constant across the Fréchet set rather than U-shaped would contradict the claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the monotonicity assumption that the paper's rule can support or reject and supplies the baseline for the monotonicity decision rule."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces potential outcomes and randomization-based inference, the foundation of the design-based model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the potential-outcome causal model and assignment-independence condition the design-based likelihood uses."},{"cited_title":"Bauld, D","cited_arxiv_id":null,"evidence_quote":"Provides one of the two empirical applications: a completely randomized payment intervention for pregnant smokers with a positive average effect."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the other empirical application: a Bernoulli-randomized organ-donation default experiment whose MLE includes 21 defiers."}],"review_version":1}