{"id":"4619703e-cce4-45f5-b7f5-36d356754751","arxiv_id":"2607.23372","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Bayesian ordered-probit modeling of joint potential outcomes yields coherent super- and finite-population inference on the probabilities that treatment is beneficial (τ) and strictly beneficial (η), far sharper than nonparametric bounds.","lead":"A Bayesian ordered-probit model gives sharp estimates of two simple probabilities: that a treatment helps or strictly helps when outcomes are only ordered categories. It is useful wherever RCTs report ratings, grades, or clinical scales and average treatment effects are hard to interpret.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The recommended sensitivity range ρ∈[0,0.5] does not actually secure inference on τ: the paper's own Table 2 shows coverage collapsing to 0.157 when true and assumed ρ differ inside that range, and the 0.5 threshold is calibrated on a single simulation DGP.","rationale":"The reader correctly located the root issue — ρ is unidentified and everything is conditional on it — and rated the paper CONDITIONAL with high confidence, which I endorse. I rate agreement \"partial\" rather than \"agree\" because the reader frames the weakness as an inherent, well-handled assumption (\"mostly well-handled... keep ρ-sensitivity front-and-center\"), whereas the sharper problem is that the specific remedy the paper offers (report ρ∈[0,0.5]) is shown by the paper's own Table 2 to be insufficient for τ even when the truth lies inside that range (coverage 0.157–0.344 off-diagonal), and the 0.5 cutoff is extrapolated from one simulation configuration despite Proposition 3 implying configuration-dependent sensitivity structure for K≥4. So the gap is not just \"ρ must be assumed\" but \"the proposed mitigation does not deliver what the abstract-level claim promises for τ.\"\n\nWhy UNCHANGED rather than a downgrade: the paper is transparent about all of this — Table 2 is published in full, the conditional-on-ρ nature of inference is stated, and the theoretical section (Theorem 1, Propositions 1–3) is a genuine contribution to understanding the sensitivity. The math I spot-checked is sound (Plackett-based derivatives, the exchangeability argument in Proposition 2, the AM–GM step). This is a solid methodology paper whose practical guidance needs one more honest step: either an explicitly conditional framing of all τ claims, or a robust aggregation rule across the sensitivity grid, plus a check that the [0,0.5] recommendation transfers across DGPs. Those are conditions, not grounds for rejection — consistent with the reader's CONDITIONAL. The misspecification concern (non-probit DGPs) the reader raised remains valid but secondary, since ρ-misspecification alone already degrades τ coverage within the paper's own framework.","tokens_in":23269,"tokens_out":2678,"duration_ms":21433,"concrete_test":"Run a \"blind analyst\" coverage study: simulate the K=5 ordered probit with true ρ ∈ {0, 0.25, 0.5} (unknown to the analyst) across several (βW, α, µ0) configurations, e.g., βW ∈ {−1.0, −0.6, −0.2} and re-spaced cutpoints. The analyst follows the paper's recommendation, reporting the τ posterior grid over assumed ρ ∈ {0, 0.1, 0.3, 0.5}. Record (a) minimum coverage of any single assumed-ρ interval when the truth is elsewhere in [0,0.5], and (b) whether the \"marked separation at ρ>0.5\" in the Figure-2 derivative plot survives across configurations. If (a) stays well below 0.95 or the threshold in (b) moves, the practical recommendation needs reformulation (e.g., a union/robust interval or explicit conditional-coverage language).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that the Bayesian ordered-probit approach \"overcomes the identifiability limitations\" and yields \"substantially sharper... practically relevant\" inference than Lu et al.'s bounds. Formally this is fine — with (θ, ρ) fixed the estimands are functions of the model. The soft spot is not that ρ must be assumed (the reader flagged this and the paper is honest about it), but that the paper's proposed practical remedy is weaker than the paper presents it, in two concrete ways.\n\n(1) Table 2's own numbers undercut the claim that reporting ρ∈[0,0.5] gives \"satisfactory coverage.\" Coverage is only high near the diagonal. With true ρ=0.5 — a value inside the recommended range — assuming ρ=0 gives τ coverage 0.157 and assuming ρ=0.1 gives 0.344; with true ρ=0.3, assuming ρ=0 gives 0.835. Since the analyst cannot know the true ρ, \"report results over ρ∈[0,0.5]\" does not deliver calibrated inference for τ even when the truth lies in the recommended range. The τ intervals are sharp only conditionally on a guess that the paper's simulation shows can be badly wrong, while the abstract-level claim of \"precise and practically relevant assessments\" reads unconditionally.\n\n(2) The 0.5 threshold itself is DGP-specific. The text derives it from Figure 2, computed for one configuration (K=5, µ0=0, βW=−0.6, α=(−∞,−2,−1,0,1,∞)): \"Figure 2 reveals a marked separation between the curves as ρ exceeds 0.5.\" But Theorem 1 and Proposition 3 show the derivative roots and hence the sensitivity pattern depend on α, µ0, βW, and K; for K≥4 multiple roots may exist. Nothing in the paper shows the separation point stays near 0.5 across plausible configurations, so the general recommendation ρ∈[0,0.5] is extrapolated from one point in parameter space.\n\nThese do not invalidate the method; they mean the central applied claim for τ inherits uncontrolled error that the proposed sensitivity analysis does not bound. The η results are genuinely more robust (flat across ρ in Table 2), so the","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript studies randomized experiments with ordinal outcomes and the estimands τ=Pr{Y(1)≤Y(0)} and η=Pr{Y(1)<Y(0)}, which are not identified from the observed marginal distributions. The authors model the two potential outcomes through a bivariate normal latent-variable ordered probit model, with the cross-potential-outcome latent correlation ρ fixed rather than estimated. They derive limiting, sign, and root-count results describing how the super-population estimands vary with ρ (Propositions 1–3 and Theorem 1); give Gibbs-sampling procedures for super-population posterior inference and finite-population imputation; compare the resulting intervals with the sharp nonparametric bounds of Lu et al. in simulations; and analyze a randomized scalp-health experiment. A sensitivity analysis over ρ leads to a proposed practical reporting range of ρ∈[0,0.5].","tokens_in":23660,"tokens_out":11092,"duration_ms":93165,"significance":"The problem is important: τ and η are substantially more interpretable for ordinal outcomes than an average effect, while their sharp nonparametric bounds are often too wide for decisions. If presented with appropriate conditioning, the proposed framework would be a useful model-based alternative. Notable strengths are the explicit treatment of both super- and finite-population inference, closed-form Gibbs updates and counterfactual imputation, analytic derivative results with detailed proofs, a transparent off-diagonal sensitivity table, and an application with covariate adjustment. The sensitivity table is especially valuable because it makes the cost of misspecifying ρ directly visible. The contribution is therefore potentially useful, but its practical force is conditional on the ordered-probit model and an assumed ρ; the present manuscript sometimes presents those conditional conclusions as if they were unconditional.","major_comments":[{"comment":"The claim that Table 2 shows “satisfactory coverage for ρ∈[0,0.5]” is not supported if interpreted as robustness across that range. High τ coverage occurs mainly on the diagonal. For example, when the true ρ=0.5—inside the recommended range—assuming ρ=0 gives coverage 0.157 and assuming ρ=0.1 gives 0.344; when true ρ=0 and assumed ρ=0.5, coverage is 0.168; true ρ=0.3 and assumed ρ=0 gives 0.835. Reporting several ρ-conditional intervals therefore does not by itself provide calibrated inference for τ. Please state this strictly as conditional sensitivity, separate the relatively stable η results from the highly sensitive τ results, and, if a range summary is recommended, define a union or model-averaged procedure and evaluate its coverage under an explicit design or prior over ρ.","section":"§5, Table 2 and following paragraph"},{"comment":"The proposed upper cutoff ρ=0.5 is inferred from one simulation configuration (K=5, μ0=0, βW=−0.6, and α=(−∞,−2,−1,0,1,∞)). Theorem 1 and Proposition 3 show that the derivative roots and sensitivity pattern depend on K, the cutpoints, μ0, and ρ, so a universal threshold at 0.5 does not follow. The statement that beyond 0.5 the unit-level correlation “begins to overshadow the treatment effect” should either be removed, framed as behavior in this particular DGP, or supported by a broad simulation grid or a quantitative theorem. Figure 2 should also identify the plotted ρ values and the numerical separation criterion.","section":"§5, Figure 2 and the recommended range"},{"comment":"The phrases “overcomes the identifiability limitations” and “provides precise and practically relevant assessments” overstate what is established. The method does not identify ρ or remove the nonidentifiability of τ and η; it replaces nonidentifiability with an ordered-probit assumption and a fixed ρ. Table 1 uses the correctly specified DGP with the true ρ=0.7 known exactly, while Table 2 shows that τ calibration can fail sharply under misspecification. The central claims should be revised to “model-based inference conditional on ρ and the assumed latent model,” with an explicit warning that nominal credibility/coverage is not unconditional over unknown ρ or model misspecification. A misspecified-latent-model simulation would be valuable, although tempering the claims is the essential revision.","section":"Abstract, §1, §5, and §7"}],"minor_comments":[{"comment":"The table reports SP truths, but the text says the Bayesian posterior means, intervals, and coverages are computed in the FP setting over 1,000 treatment assignments. Please make the coverage target explicit and avoid implying that this design validates SP posterior coverage. If SP calibration is claimed, a repeated-population simulation should be added.","section":"§5, Table 1"},{"comment":"The set T is described as a subset of R^{K+1}, although α0 and αK are fixed and only K−1 cutpoints are free. The uniform prior over the unbounded ordered polytope is improper; please give a posterior-propriety argument or cite an applicable result.","section":"§4, prior specification"},{"comment":"τsp(x~) is written as conditional on (θ,Z), but the displayed expression is calculated from pklsp(x~) and does not depend on the observed latent vector Z. Conditioning on θ, ρ, and x~ would be clearer.","section":"§4.1, Eq. (22)"},{"comment":"Please state whether “odd number of roots” means distinct roots or roots counted with multiplicity. Opposite limiting signs imply an odd total multiplicity under the present regularity, but not necessarily an odd number of distinct roots when tangencies occur.","section":"Proposition 3 and Supplement S1"},{"comment":"Because the reported posterior means and interval endpoints are medians across 1,000 assignments, the displayed lower and upper endpoints need not come from the same assignment-specific interval. Please state this explicitly and, if space permits, report Monte Carlo uncertainty.","section":"§5, Table 1"},{"comment":"The assertion that negative correlations are “generally implausible in practice” should be supported or softened. Negative dependence between potential outcomes may be unusual in some applications but is not logically impossible.","section":"§5, Table 2"},{"comment":"Clarify whether the recoded four-level baseline score enters the Bayesian linear predictor as a numeric category code. If so, this imposes equal spacing on an ordinal covariate, in tension with the paper’s motivation; category indicators or the original baseline score would avoid this issue.","section":"§6"},{"comment":"Minor corrections include “Intermediate Mean Value Theorem” → “Intermediate Value Theorem,” “SUTV A” → “SUTVA,” and “We proceed the statistical analyses” → “We proceed with the statistical analyses.” A code-availability statement, with confidential data replaced by scripts or simulated examples, would improve reproducibility.","section":"Presentation"}],"recommendation":"major_revision","confidential_remarks":"The methodology and derivations appear broadly competent and suitable for the journal. The principal barrier is not the algebra or Gibbs construction, but the calibration and generality currently attached to the proposed ρ∈[0,0.5] reporting rule."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a competent, usable extension—not a rehash—of Lu et al. (2018) and Volfovsky et al. (2015). It gives a full super- and finite-population Gibbs pipeline for τ and η under ordered probit, plus actual theory (Props 1–3, Theorem 1) on how those estimands move with the unidentified association ρ. That theory and the dual SP/FP recipe are the new pieces; the rest is standard Albert–Chib plus careful imputation.\n\nWhat it does well: the Gibbs full conditionals and FP imputation are cleanly written. Simulations under the correct DGP put posterior means near finite-population truth with ~95% coverage when ρ is known, and the intervals are dramatically tighter than Lu’s sharp bounds—which is the practical point. The scalp RCT is a real applied illustration, not a toy. They are honest that ρ is not identified and run sensitivity rather than pretending to learn it. The η results are genuinely flat across ρ in their tables; that part of the story holds up.\n\nSoft spots, in proportion: the stress-test note is right on the applied claim for τ. Table 2 shows coverage for τ collapsing well inside the recommended ρ∈[0,0.5] band when assumed ρ ≠ true ρ (e.g. true 0.5, assumed 0 → coverage 0.157). So “report over [0,0.5]” does not deliver calibrated inference for τ; it only delivers a menu of conditional answers. The 0.5 cutoff itself is read off one figure for one (K, α, βW) configuration, while their own theorem says the root structure depends on those quantities and can be non-unique for K≥4. That is a real gap between the abstract’s “precise and practically relevant” language and what the sensitivity analysis bounds. Minor relative to that: no non-probit misspecification study, no public code/data, and the usual free knobs (priors, MCMC length, pseudocount λ).\n\nNone of that makes the method incoherent. Under a stated (θ, ρ) the estimands are well-defined functions of the model, the math in the supplement is coherent, and the citation pattern is appropriate. This is for people who already want τ/η in ordinal RCTs and are willing to treat ρ as a sensitivity parameter with eyes open—especially if they care more about η than τ. It deserves a serious referee, not a desk reject. I would engage: read the theory section, use the pipeline with ρ-sensitivity front and center, and push authors to drop the over-strong reading of the [0,0.5] band.","headline":"Solid SP/FP Bayesian recipe for Lu’s τ/η with real theory on ρ; the ρ∈[0,0.5] “fix” oversells what Table 2 actually shows for τ.","tokens_in":24636,"tokens_out":666,"would_cite":false,"duration_ms":21356,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J12","62K99"],"pacs":[],"model":"grok-4.5","headline":"A Bayesian ordered-probit model turns two non-identifiable ordinal causal probabilities into sharp, decision-ready posteriors that beat wide nonparametric bounds.","keywords":["Bayesian causal inference","ordinal potential outcomes","super and finite population inference","randomized experiments","sensitivity analysis","ordered probit","treatment effect"],"falsifier":"Generate data from the stated ordered-probit model with known τ and η and correctly specified ρ; if the Bayesian 95% credible intervals miss the true finite-population values far more often than 5%, or are no narrower in practice than the nonparametric bounds, the sharpness-and-calibration claim fails.","tokens_in":24209,"feed_emoji":"📊","tokens_out":937,"duration_ms":46508,"temperature":0.7,"pith_summary":"Many trials measure ordered outcomes—pain scores, quality ratings, scalp health—where averages are hard to interpret because the category numbers are not equal steps. This paper targets two plain probabilities: that treatment is at least as good as control, and that it is strictly better. Those probabilities depend on the joint distribution of potential outcomes, which ordinary data do not identify, so existing work either assumes independence or reports bounds that are often too wide to decide anything. The authors place a continuous latent layer under the ordered categories, couple the two potential outcomes with an ordered probit model, and use Bayesian Gibbs sampling to get coherent super-population and finite-population posteriors. Simulations and a scalp-health randomized experiment show the credible intervals are much tighter than the sharp nonparametric bounds while still covering the truth when the model is right, with a sensitivity analysis for the unknown unit-level association.","feed_headline":"Bayesian model sharpens causal odds for ordered outcomes","feed_subtitle":"Ordered-probit posteriors beat wide nonparametric bounds on whether treatment helps.","key_machinery":"The ordered probit latent model: continuous potential outcomes Z(w) with treatment effect in the mean, fixed unit-level correlation ρ, and ordered cutpoints that map Z to the observed ordinal Y; Gibbs sampling either evaluates the bivariate-normal joint probabilities (super-population) or imputes missing potential outcomes (finite-population).","core_discovery":"Modeling the joint distribution of ordinal potential outcomes with a Bayesian ordered probit latent-variable model overcomes the non-identifiability of τ = Pr(Y(1) ≤ Y(0)) and η = Pr(Y(1) < Y(0)) and yields substantially sharper, coherent super-population and finite-population posterior inference than the sharp nonparametric bounds, as shown in simulations across sample sizes and category counts and in a human scalp-health trial.","pith_inferences":["The recommended ρ ≤ 0.5 band is a modeling convention; checking Gaussian versus other copula associations would test how much the sharpness claim depends on the bivariate-normal link.","Existing consumer and clinical CREs that already store ordinal endpoints could be re-analyzed for τ and η without new data collection.","Substituting ordered logit for ordered probit is an immediate robustness check the paper leaves open and would likely preserve the qualitative gain over bounds.","Covariate adjustment through the latent mean keeps the estimands one-dimensional, offering a cleaner summary than covariate-adjusted bound averages when strata are sparse."],"forward_implications":["Analysts can report one probability that treatment helps (or strictly helps) instead of K-dimensional distributional effects or bounds too wide to use.","The same model delivers both super-population and finite-population versions of τ and η from one posterior.","Any function of (τ, η)—including the relative-effect measure γ—receives automatic posterior inference.","Practical reporting can focus on ρ in [0, 0.5]; larger assumed association begins to dominate the treatment signal.","The framework extends naturally to longitudinal ordinal outcomes and to observational studies."],"fun_headline_variants":["Bayesian ordered probit sharpens ordinal causal benefit odds","Latent-variable model beats nonparametric bounds on treatment help","Joint ordered-probit posteriors identify τ and η precisely","Bayesian framework yields sharper finite-population ordinal effects","Ordered probit overcomes non-identifiability in causal odds"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The correlation between a unit’s two latent potential outcomes cannot be learned from the data and must be fixed by the analyst, so every reported probability is conditional on that choice and on the ordered-probit form being correct.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian ordered probit sharpens ordinal causal benefit odds","Latent-variable model beats nonparametric bounds on treatment help","Joint ordered-probit posteriors identify τ and η precisely","Bayesian framework yields sharper finite-population ordinal effects","Ordered probit overcomes non-identifiability in causal odds"]},"model":"grok-4.5","effort":"low","cost_usd":0.004372,"raw_usage":{"total_tokens":1223,"prompt_tokens":687,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":43724000,"prompt_tokens_details":{"text_tokens":687,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":470,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":687,"tokens_out":66,"duration_ms":8419,"temperature":1.0,"reasoning_tokens":470,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T23:59:12.799404+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate data from the stated ordered-probit model with known τ and η and correctly specified ρ; if the Bayesian 95% credible intervals miss the true finite-population values far more often than 5%, or are no narrower in practice than the nonparametric bounds, the sharpness-and-calibration claim fails.","supporting_citations":[],"review_version":1}