REVIEW 4 major objections 4 minor 12 references
A Model of a Randomized Experiment with an Application to the PROWESS Clinical Trial
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Reanalysis of the PROWESS sepsis trial estimates the drug killed two participants for every three it saved.
desk verdict A clearly written paper whose headline claim is an artifact of an arbitrary least-squares weighting, not a measured quantity: the saved/killed split is unidentified from the aggregate trial data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key object is the 4x4 matrix N(i,j) assigning each potential outcome type i to observed outcome group j, with eight cells forced to zero by logic. The observed data are expressed as a convolution of four independent binomial random variables, one for each type, with success probability p (the intended intervention fraction). The proposed least squares estimator minimizes a sum of squared randomization errors across all subsets of types, weighting each by the inverse of its variance; this objective, S(N|p), is what selects a specific point along the continuum of type vectors consistent with the aggregate data.
What would settle it
Take the reported estimates (t(2)=308, t(3)=205) and form an alternative vector with t(2)=300, t(3)=197, adjusting t(1) and t(4) to keep the expected cell counts identical; both vectors predict the same observed data, so no estimator using only these counts can distinguish them.
Extended reading notes
Core claim
The paper's central claim is that the numbers of trial participants in each of four potential outcome types—live regardless, saved, killed, die regardless—can be estimated from the experiment's aggregate outcome counts and the intended randomization fraction, without any individual-level covariates or assumptions beyond random assignment. Within the PROWESS trial, it finds 964 participants would live regardless, 308 would be saved, 205 would be killed, and 213 would die regardless, implying a ratio of two killed per three saved and 103 participants actually killed inside the trial. The author presents this as a decomposition of the reduced form, which itself only identifies the net difference between saved and killed.
Load-bearing premise
The estimator's objective function is assumed to identify the true saved/killed split, but the observed aggregate counts only fix sums such as the number saved plus the number who die regardless, so many different splits fit the data equally well; the specific 2:3 ratio is a consequence of the chosen weighting scheme, not of the data alone.
Editorial extensions
If this is right
- If the PROWESS estimates are correct, the trial's net benefit of about 6 percentage points masks a gross harm of 12% killed and a gross benefit of 18% saved.
- The method provides a template for decomposing reduced-form effects in other randomized trials into saved and killed counts, which could inform risk-benefit assessments.
- The finding implies that even a trial that supports regulatory approval can cause substantial harm to a minority of participants, raising ethical questions about informed consent and monitoring.
- The paper's estimates are consistent with the subsequent voluntary withdrawal of the drug in 2011 after a confirmatory trial failed to show survival benefit.
Reading between the lines
- Because the observed data identify only the net difference between saved and killed, the specific 2:3 ratio depends on the arbitrary weighting of randomization errors; a different defensible weighting could shift the split while leaving the reduced form unchanged.
- The method's reliance on aggregate counts means it cannot distinguish between participants killed by the drug's mechanism (e.g., bleeding) and those who would have died of sepsis anyway but were classified as killed by the potential-outcome definition; the interpretation of 'killed' is definitional, not causal at the individual level.
- The approach could be extended to instrumental-variable settings or to trials with non-binary outcomes by generalizing the type matrix and the binomial convolution.
- A bootstrap or sensitivity analysis over different weighting schemes could quantify how much of the estimated ratio is driven by the objective function rather than the data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a model of a two-arm randomized experiment with a binary outcome, classifying participants into four potential-outcome types: would live regardless, would be saved, would be killed, and would die regardless. It derives an expression for the probability of the observed cell counts as a function of the type counts and the randomization probability, then proposes a maximum likelihood estimator and a computationally tractable least squares estimator. Applying the least squares estimator to aggregate data from the PROWESS trial (cell counts g = (210, 640, 259, 581), p = 1/2), the paper reports that 18% of participants would be saved, 12% would be killed, that the intervention killed two participants for every three saved, and that 103 trial participants were killed by the intervention. A Monte Carlo simulation calibrates the true type counts to these point estimates and reports bias and RMSE.
Significance. Were the estimates identified by the data, the ethical and clinical implications would be substantial, because the paper claims to quantify the number of trial participants harmed by a life-saving intervention. The paper is also transparent about the computational difficulty of the maximum likelihood problem and clearly explains the potential-outcome taxonomy. However, the central quantitative claim is not supported: the observed data identify only the reduced-form difference between saved and killed counts, not the split between them, so the reported 103 killed and the 2:3 ratio are artifacts of an arbitrary objective function rather than evidence. The Monte Carlo exercise cannot repair this because it assumes the disputed point estimates as the truth.
major comments (4)
- [Section 2, Eqs. (1)-(3)] The likelihood expression is invariant to the transformation t -> t + a*(1, -1, -1, 1), which preserves all four sums t3+t4, t1+t2, t2+t4, and t1+t3 that enter the data probabilities. Consequently, the observed cell counts identify only the marginal mortality rates in the two arms, and the split between participants who would be saved and those who would be killed is unidentified. The maximum likelihood estimator is therefore set-valued over a continuum of type-count vectors, and the likelihood surface is flat in the saved/killed dimension.
- [Section 3.2, objective S(N|p)] The least squares estimator minimizes a weighted sum of squared randomization errors over the unobserved matrix N(i,j), subject only to the column-sum constraints that reproduce the observed cell counts. Because those constraints leave the row split (equivalently, the saved/killed decomposition) free, the minimum of S(N|p) is determined by the particular weighting scheme rather than by features of the data. The reported t(2)=308 and t(3)=205 are therefore selected by the objective function, not identified by the experiment.
- [Section 4.2, Figure 2] The bootstrapped standard error for the estimated number killed, t(3)=205, is reported as 132, so the estimate is not statistically distinguishable from zero at conventional levels. The claim that the intervention killed 103 participants within the trial relies on an intermediate cell estimate n(3,1)=103 with standard error 66, which is also not statistically significant. The paper's headline ratio and absolute killed count are thus unsupported by its own reported precision.
- [Section 5, Monte Carlo design] The Monte Carlo simulation sets the true potential-outcome type vector t equal to the paper's own point estimates from PROWESS and then evaluates the estimator's bias and RMSE. This exercise can at most quantify finite-sample noise around a chosen oracle value; it cannot establish that the observed PROWESS data identify the saved/killed split, because the simulation never varies the split along the unidentified dimension. The reported mean bias of about 42 participants and RMSE of about 124 participants are therefore not evidence for the validity of the point estimates.
minor comments (4)
- [Introduction, paragraph after the abstract results] There is a duplicated word: 'about half of those participants were were randomized into the control group' should read 'were randomized'.
- [Section 2, paragraph before Eq. (1)] The phrase 'effectively a separate randomized experiment within each potential outcome outcome type i' contains a duplicated word 'outcome' and should be corrected.
- [Title page note] The manuscript states it has been combined with another paper and superseded by a later arXiv posting. If this version is still under consideration, the relationship to the superseding paper should be clarified in the submission letter or a revision.
- [Section 4.1, first paragraph] The reduced form is reported as '6 percentage points', while the precise calculation 210/850 - 259/840 equals approximately -0.0602; stating the exact value in the text would avoid ambiguity.
Circularity Check
Saved/killed split is unidentified; reported 2:3 ratio is imposed by the least-squares objective, and Monte Carlo validates the estimator against its own estimates.
-
self definitional
[Section 3.1 (flat likelihood); Eq. (3) in Section 2; estimates in Section 4.2]
"It is possible that this problem could be solved numerically with another software package via a grid search, but I have not been able to find a solution through such an approach because the size of the grid is large for experiments of reasonable size, and evaluation of the objective function produces numbers that are the same within machine precision."
The four observed cells are multinomial with probabilities proportional to p(t3+t4), p(t1+t2), (1-p)(t2+t4), and (1-p)(t1+t3), so the likelihood is invariant under t -> t + a(1,-1,-1,1): t(2) and t(3) are not separately identified. The flat likelihood that the paper reports over the grid is the computational signature of this non-identification. The reported t(2)=308 and t(3)=205 are the minimizer of the author's chosen weighted least-squares objective S(N|p), not a feature of the data. Thus the headline 18%/12% split and the 'two killed for every three saved' ratio are supplied by the construction of the estimator, not derived from the experiment.
-
self definitional
[Section 5.1 (Monte Carlo design); Section 5.2 (Monte Carlo results)]
"In each simulated experiment m, I set the true vector of the number of participants of each potential outcome type t to the estimated vector from the PROWESS trial."
The Monte Carlo defines the estimator's own point estimates as the true potential outcome counts, then asks whether the estimator recovers them. Because t(2) and t(3) are unidentified by the likelihood, this exercise cannot establish that the saved/killed estimates are correct; it only checks internal consistency of the optimizer under the model. The reported small bias and RMSE are therefore self-referential: the simulation assumes the disputed t(2)=308 and t(3)=205 are true and then recovers them, so it provides no independent support for the central 2:3 claim.
full rationale
The paper's reduced-form result (a 6 percentage point net mortality reduction) is a legitimate data summary based on the observed cell counts. The circularity enters at the decomposition into saved and killed participants. In Eq. (3), the likelihood depends on t only through sums such as t(2)+t(4) and t(1)+t(3), so the observed intervention/control mortality counts cannot separate participants who would be saved from participants who would be killed. The paper itself notes that evaluating the likelihood gives values identical within machine precision, the expected signature of a flat/unidentified objective. The least-squares estimator then selects a point on that flat manifold by minimizing a weighted sum of squared randomization errors, S(N|p), with weights chosen inversely proportional to variances; the reported 57%/18%/12%/13% splits and the intermediate estimate of 103 in-trial killed are outputs of that arbitrary weighting, not of the data. The Monte Carlo exercise is also self-referential: it sets the true t equal to the paper's own point estimates, so it cannot validate identification of t(2) and t(3); it only verifies that the optimizer can recover its own inputs. No load-bearing self-citation is involved. Because the central ethical claim—two participants killed for every three saved—reduces to the estimator's construction rather than to evidence in the trial data, the circularity score is 8.
Assumptions & free parameters
assumptions (4)
- domain assumption Random assignment to intervention is independent across participants with the same probability p.
- domain assumption Each participant has a deterministic potential outcome type among exactly four types, with no noncompliance or interference.
- ad hoc to paper The least squares objective S(N) is an appropriate loss for selecting the true N.
- ad hoc to paper The Monte Carlo true t equals the PROWESS point estimates.
Cite this review
Pith. "Pith review of A Model of a Randomized Experiment with an Application to the PROWESS Clinical Trial." pith.science (2026). https://pith.science/paper/OX3UWZRH
@misc{pith2026190805810,
author = {Pith},
title = {Pith review of: A Model of a Randomized Experiment with an Application to the PROWESS Clinical Trial},
year = {2026},
howpublished = {\url{https://pith.science/paper/OX3UWZRH}},
note = {Machine review of arXiv:1908.05810}
}
read the original abstract
I develop a model of a randomized experiment with a binary intervention and a binary outcome. Potential outcomes in the intervention and control groups give rise to four types of participants. Fixing ideas such that the outcome is mortality, some participants would live regardless, others would be saved, others would be killed, and others would die regardless. These potential outcome types are not observable. However, I use the model to develop estimators of the number of participants of each type. The model relies on the randomization within the experiment and on deductive reasoning. I apply the model to an important clinical trial, the PROWESS trial, and I perform a Monte Carlo simulation calibrated to estimates from the trial. The reduced form from the trial shows a reduction in mortality, which provided a rationale for FDA approval. However, I find that the intervention killed two participants for every three it saved.
Reference graph
Works this paper leans on
-
[1]
Bernard, G. R., J.-L. Vincent, P.-F. Laterre, S. P. LaRosa, J.-F. Dhainaut, A. Lopez-Rodriguez, J. S. Steingrub, G. E. Garber, J. D. Helterbrand, E. W. Ely, and C. J. Fisher (2001). Efficacy and safety of recombinant human activated protein c for severe sepsis. New England Journal of Medicine\/ 344\/ (10), 699--709. PMID: 11236773
work page 2001
-
[2]
Foot, P. (1967). The problem of abortion and the doctrine of double effect. Oxford Review\/ 5 , 5--15
work page 1967
-
[3]
Holland, P. W. (1986). Statistics and causal inference. Journal of the American statistical Association\/ 81\/ (396), 945--960
1986
-
[4]
Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology\/ 66\/ (5), 688
1974
-
[5]
Rubin, D. B. (1977). Assignment to treatment group on the basis of a covariate. Journal of Educational and Behavioral statistics\/ 2\/ (1), 1--26
work page 1977
-
[6]
Sahinidis, N. V. (2018). BARON 18.8.23: Global Optimization of Mixed-Integer Nonlinear Programs, User's Manual
work page 2018
-
[7]
Siegel, J. P. (2002). Assessing the use of activated protein c in the treatment of severe sepsis. The New England journal of medicine\/ 347\/ (13), 1030--1034
work page 2002
-
[8]
The Mathworks , Inc. (2016). Matlab, V ersion r2016a. Natick, Massachusetts, United States
work page 2016
Show all 12 references
-
[9]
Thomson, J. J. (1985). The trolley problem. Yale Law Journal\/ 94\/ (1395), 1395--1415
1985
-
[10]
FDA drug safety communication: voluntary market withdrawal of xigris due to failure to show a survival benefit
US Food and Drug Administration and others (2011). FDA drug safety communication: voluntary market withdrawal of xigris due to failure to show a survival benefit. US Food and Drug Administration, Washington, DC\/
2011
-
[11]
Warren, H. S., A. F. Suffredini, P. Q. Eichacker, and R. S. Munford (2002). Risks and benefits of activated protein c treatment for severe sepsis. The New England journal of medicine\/ 347\/ (13), 1027--1030
2002
-
[12]
Wolfram Research , Inc. (2018). Mathematica, V ersion 11.3. Champaign, IL
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.