Pith. sign in

REVIEW 3 major objections 8 minor 34 references

A Bayesian ordered-probit model turns two non-identifiable ordinal causal probabilities into sharp, decision-ready posteriors that beat wide nonparametric bounds.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 23:59 UTC pith:62VE22MM

load-bearing objection Solid SP/FP Bayesian recipe for Lu’s τ/η with real theory on ρ; the ρ∈[0,0.5] “fix” oversells what Table 2 actually shows for τ. the 3 major comments →

arxiv 2607.23372 v1 pith:62VE22MM submitted 2026-07-25 stat.ME math.STstat.TH

Causal Inference of Ordinal Outcomes: A Bayesian Solution

classification stat.ME math.STstat.TH MSC 62F1562J1262K99
keywords Bayesian causal inferenceordinal potential outcomessuper and finite population inferencerandomized experimentssensitivity analysisordered probittreatment effect
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Many trials measure ordered outcomes—pain scores, quality ratings, scalp health—where averages are hard to interpret because the category numbers are not equal steps. This paper targets two plain probabilities: that treatment is at least as good as control, and that it is strictly better. Those probabilities depend on the joint distribution of potential outcomes, which ordinary data do not identify, so existing work either assumes independence or reports bounds that are often too wide to decide anything. The authors place a continuous latent layer under the ordered categories, couple the two potential outcomes with an ordered probit model, and use Bayesian Gibbs sampling to get coherent super-population and finite-population posteriors. Simulations and a scalp-health randomized experiment show the credible intervals are much tighter than the sharp nonparametric bounds while still covering the truth when the model is right, with a sensitivity analysis for the unknown unit-level association.

Core claim

Modeling the joint distribution of ordinal potential outcomes with a Bayesian ordered probit latent-variable model overcomes the non-identifiability of τ = Pr(Y(1) ≤ Y(0)) and η = Pr(Y(1) < Y(0)) and yields substantially sharper, coherent super-population and finite-population posterior inference than the sharp nonparametric bounds, as shown in simulations across sample sizes and category counts and in a human scalp-health trial.

What carries the argument

The ordered probit latent model: continuous potential outcomes Z(w) with treatment effect in the mean, fixed unit-level correlation ρ, and ordered cutpoints that map Z to the observed ordinal Y; Gibbs sampling either evaluates the bivariate-normal joint probabilities (super-population) or imputes missing potential outcomes (finite-population).

Load-bearing premise

The correlation between a unit’s two latent potential outcomes cannot be learned from the data and must be fixed by the analyst, so every reported probability is conditional on that choice and on the ordered-probit form being correct.

What would settle it

Generate data from the stated ordered-probit model with known τ and η and correctly specified ρ; if the Bayesian 95% credible intervals miss the true finite-population values far more often than 5%, or are no narrower in practice than the nonparametric bounds, the sharpness-and-calibration claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Analysts can report one probability that treatment helps (or strictly helps) instead of K-dimensional distributional effects or bounds too wide to use.
  • The same model delivers both super-population and finite-population versions of τ and η from one posterior.
  • Any function of (τ, η)—including the relative-effect measure γ—receives automatic posterior inference.
  • Practical reporting can focus on ρ in [0, 0.5]; larger assumed association begins to dominate the treatment signal.
  • The framework extends naturally to longitudinal ordinal outcomes and to observational studies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The recommended ρ ≤ 0.5 band is a modeling convention; checking Gaussian versus other copula associations would test how much the sharpness claim depends on the bivariate-normal link.
  • Existing consumer and clinical CREs that already store ordinal endpoints could be re-analyzed for τ and η without new data collection.
  • Substituting ordered logit for ordered probit is an immediate robustness check the paper leaves open and would likely preserve the qualitative gain over bounds.
  • Covariate adjustment through the latent mean keeps the estimands one-dimensional, offering a cleaner summary than covariate-adjusted bound averages when strata are sparse.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The manuscript studies randomized experiments with ordinal outcomes and the estimands τ=Pr{Y(1)≤Y(0)} and η=Pr{Y(1)<Y(0)}, which are not identified from the observed marginal distributions. The authors model the two potential outcomes through a bivariate normal latent-variable ordered probit model, with the cross-potential-outcome latent correlation ρ fixed rather than estimated. They derive limiting, sign, and root-count results describing how the super-population estimands vary with ρ (Propositions 1–3 and Theorem 1); give Gibbs-sampling procedures for super-population posterior inference and finite-population imputation; compare the resulting intervals with the sharp nonparametric bounds of Lu et al. in simulations; and analyze a randomized scalp-health experiment. A sensitivity analysis over ρ leads to a proposed practical reporting range of ρ∈[0,0.5].

Significance. The problem is important: τ and η are substantially more interpretable for ordinal outcomes than an average effect, while their sharp nonparametric bounds are often too wide for decisions. If presented with appropriate conditioning, the proposed framework would be a useful model-based alternative. Notable strengths are the explicit treatment of both super- and finite-population inference, closed-form Gibbs updates and counterfactual imputation, analytic derivative results with detailed proofs, a transparent off-diagonal sensitivity table, and an application with covariate adjustment. The sensitivity table is especially valuable because it makes the cost of misspecifying ρ directly visible. The contribution is therefore potentially useful, but its practical force is conditional on the ordered-probit model and an assumed ρ; the present manuscript sometimes presents those conditional conclusions as if they were unconditional.

major comments (3)
  1. [§5, Table 2 and following paragraph] The claim that Table 2 shows “satisfactory coverage for ρ∈[0,0.5]” is not supported if interpreted as robustness across that range. High τ coverage occurs mainly on the diagonal. For example, when the true ρ=0.5—inside the recommended range—assuming ρ=0 gives coverage 0.157 and assuming ρ=0.1 gives 0.344; when true ρ=0 and assumed ρ=0.5, coverage is 0.168; true ρ=0.3 and assumed ρ=0 gives 0.835. Reporting several ρ-conditional intervals therefore does not by itself provide calibrated inference for τ. Please state this strictly as conditional sensitivity, separate the relatively stable η results from the highly sensitive τ results, and, if a range summary is recommended, define a union or model-averaged procedure and evaluate its coverage under an explicit design or prior over ρ.
  2. [§5, Figure 2 and the recommended range] The proposed upper cutoff ρ=0.5 is inferred from one simulation configuration (K=5, μ0=0, βW=−0.6, and α=(−∞,−2,−1,0,1,∞)). Theorem 1 and Proposition 3 show that the derivative roots and sensitivity pattern depend on K, the cutpoints, μ0, and ρ, so a universal threshold at 0.5 does not follow. The statement that beyond 0.5 the unit-level correlation “begins to overshadow the treatment effect” should either be removed, framed as behavior in this particular DGP, or supported by a broad simulation grid or a quantitative theorem. Figure 2 should also identify the plotted ρ values and the numerical separation criterion.
  3. [Abstract, §1, §5, and §7] The phrases “overcomes the identifiability limitations” and “provides precise and practically relevant assessments” overstate what is established. The method does not identify ρ or remove the nonidentifiability of τ and η; it replaces nonidentifiability with an ordered-probit assumption and a fixed ρ. Table 1 uses the correctly specified DGP with the true ρ=0.7 known exactly, while Table 2 shows that τ calibration can fail sharply under misspecification. The central claims should be revised to “model-based inference conditional on ρ and the assumed latent model,” with an explicit warning that nominal credibility/coverage is not unconditional over unknown ρ or model misspecification. A misspecified-latent-model simulation would be valuable, although tempering the claims is the essential revision.
minor comments (8)
  1. [§5, Table 1] The table reports SP truths, but the text says the Bayesian posterior means, intervals, and coverages are computed in the FP setting over 1,000 treatment assignments. Please make the coverage target explicit and avoid implying that this design validates SP posterior coverage. If SP calibration is claimed, a repeated-population simulation should be added.
  2. [§4, prior specification] The set T is described as a subset of R^{K+1}, although α0 and αK are fixed and only K−1 cutpoints are free. The uniform prior over the unbounded ordered polytope is improper; please give a posterior-propriety argument or cite an applicable result.
  3. [§4.1, Eq. (22)] τsp(x~) is written as conditional on (θ,Z), but the displayed expression is calculated from pklsp(x~) and does not depend on the observed latent vector Z. Conditioning on θ, ρ, and x~ would be clearer.
  4. [Proposition 3 and Supplement S1] Please state whether “odd number of roots” means distinct roots or roots counted with multiplicity. Opposite limiting signs imply an odd total multiplicity under the present regularity, but not necessarily an odd number of distinct roots when tangencies occur.
  5. [§5, Table 1] Because the reported posterior means and interval endpoints are medians across 1,000 assignments, the displayed lower and upper endpoints need not come from the same assignment-specific interval. Please state this explicitly and, if space permits, report Monte Carlo uncertainty.
  6. [§5, Table 2] The assertion that negative correlations are “generally implausible in practice” should be supported or softened. Negative dependence between potential outcomes may be unusual in some applications but is not logically impossible.
  7. [§6] Clarify whether the recoded four-level baseline score enters the Bayesian linear predictor as a numeric category code. If so, this imposes equal spacing on an ordinal covariate, in tension with the paper’s motivation; category indicators or the original baseline score would avoid this issue.
  8. [Presentation] Minor corrections include “Intermediate Mean Value Theorem” → “Intermediate Value Theorem,” “SUTV A” → “SUTVA,” and “We proceed the statistical analyses” → “We proceed with the statistical analyses.” A code-availability statement, with confidential data replaced by scripts or simulated examples, would improve reproducibility.

Circularity Check

0 steps flagged

No significant circularity: estimands, model, and sensitivity analysis are independent of one another; self-citations supply problem setup, not the solution.

full rationale

The paper’s load-bearing chain is: (i) define τ and η from the joint of ordinal potential outcomes (Lu et al. 2018 and related work); (ii) note that these functionals are non-identified from margins alone when K≥3; (iii) impose an ordered-probit latent joint with fixed association ρ and run Bayesian imputation for super- and finite-population posteriors; (iv) derive analytic sensitivity of τsp, ηsp to ρ (Props. 1–3, Thm. 1); (v) compare posteriors to nonparametric sharp bounds and to known truth under the DGP. None of these steps reduces by construction to its inputs. τ and η are not defined from the fitted ordered-probit parameters; the model is an identifying assumption, openly conditional on ρ, with explicit sensitivity rather than a claim that ρ is learned. Simulation “accuracy” is checked against DGP truth, not against a quantity that was fitted and then relabeled as a prediction. Self-citations (Lu–Ding–Dasgupta bounds line; Dasgupta et al. on ρ; Volfovsky et al. latent framework) supply the estimands being targeted and standard background on missing potential outcomes; they do not import a uniqueness theorem or ansatz that forces the present posterior claims. The skeptic’s concerns about the recommended ρ∈[0,0.5] range and Table 2 coverage are robustness/calibration issues, not circularity. Derivation is self-contained against external benchmarks under stated assumptions.

Axiom & Free-Parameter Ledger

4 free parameters · 7 axioms · 0 invented entities

The inferential engine rests on standard causal design assumptions plus a parametric ordered-probit joint for potential outcomes. Identifiability of τ and η is purchased by fixing the latent correlation ρ and the normal/cutpoint structure; those are the load-bearing modeling inputs the reader did not get ‘for free’ from the observed margins alone.

free parameters (4)
  • ρ (latent potential-outcome correlation) = assumed; sims use 0.7; application reports 0, 0.1, 0.3, 0.5
    Unidentifiable from observed data; fixed by analyst and swept in sensitivity. All reported τ/η posteriors condition on it. Recommended operating range [0, 0.5] is a judgment call (§3, §5 Table 2).
  • Prior hyperparameters b0, B0 for β = b0=0, B0=100 in simulations
    Conjugate Normal prior mean/covariance chosen by authors (sims: b0=0, B0=100); affects finite-sample posterior (§4 Prior; §5 initials).
  • Pseudocount λ for covariate-conditional margins = 0.01
    Ad-hoc λ=0.01 used so conditional bound estimates exist when a stratum lacks treated or control units in the scalp data (§6, eq. 26).
  • MCMC length and burn-in (nM, nB) = nM=20000, nB=15000
    Operational choices that affect reported posterior means/CIs; set to 20,000 / 15,000 after Gelman–Rubin checks (§5).
axioms (7)
  • domain assumption SUTVA: no interference and no hidden versions of treatment; potential outcomes (Yi(1), Yi(0)) well-defined.
    Invoked at the start of the potential-outcomes setup (§2, Rubin 1980).
  • domain assumption Completely randomized experiment implies ignorable assignment: (Y(1), Y(0)) ⊥ W.
    Used to drop the assignment mechanism from the likelihood (§2 eq. 1; §4 after de Finetti).
  • domain assumption Identical discretizing map g for treatment and control latents (common cutpoints α).
    Required so ordinal-scale estimands τ, η are causally meaningful under the latent model (§3 before eq. 14).
  • domain assumption Latent residuals are bivariate normal with unit variances and correlation ρ (ordered probit).
    Core parametric assumption enabling joint probabilities via Φ2 (§3 eqs. 14–16, 20).
  • standard math Identification normalizations: drop intercept, fix latent scale σ=1, leave cutpoints ordered but otherwise free.
    Standard ordered-probit identification (Jackman 2009), adopted explicitly (§3).
  • domain assumption Lower ordinal categories are better; sign of βW is interpreted accordingly.
    Convention fixed to match the scalp study; reverses direction relative to much of the cited literature (Remark 3.1; §2.1).
  • standard math Independence of units and de Finetti-style i.i.d. sampling of complete data given θ for super-population inference.
    Stated at opening of §4 to justify the complete-data likelihood factorization.

pith-pipeline@v1.2.0-grok45-kimik3 · 27372 in / 4017 out tokens · 75424 ms · 2026-07-30T23:59:12.799404+00:00 · methodology

0 comments
read the original abstract

Randomized experiments with ordinal outcomes are common in many scientific applications, but conventional causal estimands such as the average treatment effect are difficult to interpret because ordinal categories lack meaningful numerical spacing. We develop a Bayesian latent variable framework for drawing coherent super population and finite population inference on two interpretable causal estimands that quantify the probabilities that treatment is beneficial and strictly beneficial. By modeling the joint distribution of potential outcomes through an ordered probit model, the proposed approach overcomes the identifiability limitations of existing methods and yields substantially sharper inference than nonparametric bounds. We also investigate the impact of the unknown association between potential outcomes and propose a sensitivity analysis to assess its influence. Simulation studies and an application to a randomized experiment on human scalp health demonstrate that the method provides precise and practically relevant assessments of treatment effectiveness.

Figures

Figures reproduced from arXiv: 2607.23372 by Pradipta Sarkar, Rituparna Dey, Tirthankar Dasgupta.

Figure 1
Figure 1. Figure 1: Traceplots of β, α1 and α2 for the 3-category model with N = 100, 250 and 500 units based on five independent MCMC chains with different parameter initializations. Each column corresponds to the parameters for a specific N, while the rows display the changes in the estimated parameters with increasing N. The Gelman Rubin Rb statistic is shown at the top right corner of each plot, recording the point estima… view at source ↗
Figure 2
Figure 2. Figure 2: Plot of τ ′ ρ (βW ) vs βW (left) and η ′ ρ (βW ) vs βW (right) for the 5-category model with parameters µ0 = 0, βW = −0.6 (dashed line) and α = (−∞, −2, −1, 0, 1, ∞). as the median of those 1, 000 posterior means. Similarly, the reported 95% credible interval limits are the medians of the combined lower and upper limits respectively. Posterior coverage is the mean across the treatment assignments. The resu… view at source ↗
Figure 3
Figure 3. Figure 3: Head zones and corresponding posterior heatmaps. [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Plot of τ ′ ρ (βW ) vs βW and η ′ ρ (βW ) vs βW for the human scalp health experiment. 31 [PITH_FULL_IMAGE:figures/full_fig_p031_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references

  1. [1]

    Estimating causal effects of treatments in randomized and nonrandomized studies , volume =

    Rubin, Donald , journal =. Estimating causal effects of treatments in randomized and nonrandomized studies , volume =

  2. [2]

    Statistical Science , volume =

    Neyman , title =. Statistical Science , volume =. 1923 , note =

  3. [3]

    Rubin , journal =

    Donald B. Rubin , journal =. Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment , urldate =

  4. [4]

    Holland , journal =

    Paul W. Holland , journal =. Statistics and Causal Inference , volume =

  5. [5]

    Rubin , journal =

    Donald B. Rubin , journal =. Bayesian Inference for Causal Effects: The Role of Randomization , urldate =

  6. [6]

    Airoldi and Donald B

    Alexander Volfovsky and Edoardo M. Airoldi and Donald B. Rubin , journal =. Causal inference for ordinal outcomes , year =

  7. [7]

    Bayesian Causal Inference: A Critical Review , journal =

    Li, Fan and Ding, Peng and Mealli, Fabrizia , year =. Bayesian Causal Inference: A Critical Review , journal =

  8. [8]

    Albert and Siddhartha Chib , journal =

    James H. Albert and Siddhartha Chib , journal =. Bayesian Analysis of Binary and Polychotomous Response Data , urldate =

  9. [9]

    Nonparametric analysis of treatment effects in ordered response models , volume =

    Boes, Stefan , journal =. Nonparametric analysis of treatment effects in ordered response models , volume =

  10. [10]

    Criteria for Surrogate end Points Based on Causal Distributions , volume =

    Ju, Chuan and Geng, Zhi , journal =. Criteria for Surrogate end Points Based on Causal Distributions , volume =

  11. [11]

    Treatment Effects on Ordinal Outcomes: Causal Estimands and Sharp Bounds , volume =

    Jiannan Lu and Peng Ding and Tirthankar Dasgupta , journal =. Treatment Effects on Ordinal Outcomes: Causal Estimands and Sharp Bounds , volume =

  12. [12]

    Sharp nonparametric bounds and randomization inference for treatment effects on an ordinal outcome , volume =

    Chiba, Yasutaka , journal =. Sharp nonparametric bounds and randomization inference for treatment effects on an ordinal outcome , volume =

  13. [13]

    Rubin , journal =

    Donald B. Rubin , journal =. Causal Inference Using Potential Outcomes: Design, Modeling, Decisions , volume =

  14. [14]

    Bayesian Inference of Causal Effects for an Ordinal Outcome in Randomized Trials , number =

    Yasutaka Chiba , journal =. Bayesian Inference of Causal Effects for an Ordinal Outcome in Randomized Trials , number =

  15. [15]

    Sharp Bounds on the Relative Treatment Effect for Ordinal Outcomes , volume =

    Lu, Jiannan and Zhang, Yunshu and Ding, Peng , journal =. Sharp Bounds on the Relative Treatment Effect for Ordinal Outcomes , volume =

  16. [16]

    Horowitz and Charles F

    Joel L. Horowitz and Charles F. Manski , journal =. Nonparametric Analysis of Randomized Experiments with Missing Covariate and Outcome Data , urldate =

  17. [17]

    Bayesian Analysis for the Social Sciences , year =

    Jackman, Simon , date-added =. Bayesian Analysis for the Social Sciences , year =

  18. [18]

    Rubin , journal =

    Andrew Gelman and Donald B. Rubin , journal =

  19. [19]

    and Rubin, Donald B

    Imbens, Guido W. and Rubin, Donald B. , year=. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction , publisher=

  20. [20]

    2008 , howpublished =

    Chatterjee, Sourav , title =. 2008 , howpublished =

  21. [21]

    , title =

    Nikulin, M.S. , title =. Encyclopedia of Mathematics , publisher =. 2001 , note =

  22. [22]

    Annals of Mathematical Statistics , year=

    On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other , author=. Annals of Mathematical Statistics , year=

  23. [23]

    Tanner and Wing Hung Wong , journal =

    Martin A. Tanner and Wing Hung Wong , journal =. The Calculation of Posterior Distributions by Data Augmentation , urldate =

  24. [24]

    Miratrix , title =

    Peng Ding and Xinran Li and Luke W. Miratrix , title =. Journal of Causal Inference , number =

  25. [25]

    Exchangeability, Correlation, and Bayes' Effect , volume =

    O'Neill, Ben , journal =. Exchangeability, Correlation, and Bayes' Effect , volume =

  26. [26]

    Patel and Nayan K Patel and Apeksha M

    Maheshvari N. Patel and Nayan K Patel and Apeksha M. Merja and Dhruvil Gajera and Arnav K. Purani and Jemini H. Pandya , journal =. Methodology Validation: Correlating Adherent Scalp Flaking Score (ASFS) With Phototrichogram for Scalp Dandruff Evaluation in Adult Subjects , volume =

  27. [27]

    Agresti, Alan , keywords =

  28. [28]

    McKelvey and William James Zavoina , journal =

    Richard D. McKelvey and William James Zavoina , journal =. A statistical model for the analysis of ordinal level dependent variables , volume =

  29. [29]

    Pillai and Donald B

    Tirthankar Dasgupta and Natesh S. Pillai and Donald B. Rubin , journal =. Causal inference from 2^K factorial designs by using potential outcomes , urldate =

  30. [30]

    R. L. Plackett , journal =. A Reduction Formula for Normal Multivariate Integrals , urldate =

  31. [31]

    Journal d’Analyse Math

    On Polya frequency functions , author=. Journal d’Analyse Math. 1951 , volume=

  32. [32]

    G. J. O. Jameson , journal =. Counting Zeros of Generalised Polynomials: Descartes' Rule of Signs and Laguerre's Extensions , urldate =

  33. [33]

    Bacon and Haruko Mizoguchi and James R

    Robert A. Bacon and Haruko Mizoguchi and James R. Schwartz , journal =. Assessing therapeutic effectiveness of scalp treatments for dandruff and seborrheic dermatitis, part 1: a reliable and relevant method based on the adherent scalp flaking score (ASFS) , volume =

  34. [34]

    Locker, Kathryn C. S. and Bacon, Robert A. and Caterino, Tamara L. and Breyfogle, Laurie and Alperet, Derrick Johnston and Sarkar, Pradipta and Piliang, Melissa and Davis, Michael G. , title =. International Journal of Cosmetic Science , volume =