Pith. sign in

REVIEW 4 major objections 3 references

Fractional weights let censored patient pairs still inform hierarchical win statistics without changing the target.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 15:56 UTC pith:Y5UBEPNB

load-bearing objection Abstract-only methods idea for fractional weighting of censoring-induced ties in hierarchical win statistics; full text is the wrong paper, so identification and efficiency claims stay unchecked. the 4 major comments →

arxiv 2605.26507 v3 pith:Y5UBEPNB submitted 2026-05-26 stat.ME

Making censored pairs count: conditional tie weighting for win statistics with composite survival endpoints

classification stat.ME MSC 62N0162G0562P10
keywords win statisticshierarchical composite endpointsright censoringconditional tie weightingrestricted-time estimandU-statisticswin rationet benefit
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

In trials that rank outcomes (for example death first, then hospitalization), a pair of patients is compared only on the next outcome when the higher-priority comparison is a genuine tie. Right censoring often leaves that higher-priority tie unconfirmed even when the lower-priority outcome is fully observed, so standard restricted win-statistic estimators discard the pair entirely. This paper replaces the missing genuine-tie indicator with its conditional probability given what was observed for that pair. The resulting estimator still targets the same restricted-time win probabilities, net benefit, and win odds, but lets partially observed pairs contribute a fractional weight when their lower-priority comparison is informative. Theory for two-sample U-statistics with estimated nuisance functions, sandwich variances, simulations, and a reanalysis of a heart-failure trial support large efficiency gains under heavier censoring and longer restriction times.

Core claim

Conditional tie weighting recovers the same restricted-time hierarchical win probabilities as existing all-or-nothing restricted win-statistic estimators while allowing pairs with censoring-induced higher-priority ties to contribute fractionally whenever the lower-priority comparison is observed and informative.

What carries the argument

Conditional tie weight: the unavailable higher-priority genuine-tie indicator is replaced by its conditional probability given the observed pairwise data, turning a hard exclusion into a fractional contribution inside two-sample U-statistics with estimated nuisance functions.

Load-bearing premise

The method works only if the conditional probability of a true higher-priority tie, given the observed pair data, is correctly identified and estimated; if that nuisance model is wrong, the fractional weights can shift the estimand rather than merely improve precision.

What would settle it

In a simulation with known restricted-time win probabilities, heavy censoring, and deliberately misspecified tie-probability nuisances, check whether the conditional-tie-weight estimator remains unbiased for the restricted-time target (or whether bias appears while variance still falls).

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under heavy censoring or long restriction horizons, hierarchical win ratio, net benefit, and win odds can be estimated with substantially smaller variance without changing the scientific target.
  • Pairs that currently contribute nothing in death-first hospitalization hierarchies can still inform treatment comparisons when hospitalization is observed.
  • Sandwich variance formulas for win ratio, net benefit, and win odds become available for the fractionally weighted U-statistics.
  • Completed trials with hierarchical composite endpoints can be reanalyzed to recover information previously discarded by all-or-nothing restricted estimators.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same conditional-probability idea may extend to hierarchies with more than two priority levels if intermediate ties can be modeled analogously.
  • If the nuisance models for tie probabilities can be made robust or doubly robust, the method could tolerate more realistic censoring and dependence structures common in multi-event survival data.
  • Efficiency gains large enough under heavy censoring may change power calculations and sample-size planning for hierarchical composite primary endpoints.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. From the abstract alone, the paper proposes conditional tie weighting for hierarchical composite survival endpoints under right censoring. Existing restricted win-statistic estimators require higher-priority genuine ties to be fully observed before lower-priority comparisons can contribute; the proposed method replaces the unobserved higher-priority genuine-tie indicator by its conditional probability given the observed pairwise data, so partially observed pairs can contribute fractionally. The abstract claims this targets the same restricted-time win probabilities as the all-or-nothing restricted estimators, develops identification and large-sample theory for two-sample U-statistics with estimated nuisances, supplies sandwich variances for win ratio, net benefit, and win odds, and reports efficiency gains plus an HF-ACTION reanalysis. The body text supplied with the submission is not this manuscript: it is an unrelated paper (SIKA-GP on sparse inducing-kernel approximations for Gaussian processes).

Significance. If the identification claim holds under standard censoring and nuisance models, conditional tie weighting would be a useful methodological contribution: it would preserve a clinically interpretable restricted-time estimand while recovering information that current restricted win statistics discard under heavy censoring or long restriction horizons. That would matter for cardiovascular and other trials that use death-first hierarchical composites. The claimed sandwich variances and U-statistic theory with estimated nuisances would also be practically valuable. These strengths cannot be credited on the present file, because the proofs, assumptions, simulations, and HF-ACTION analysis are not present in the supplied full text.

major comments (4)
  1. Manuscript mismatch: the title, abstract, and arXiv id (2605.26507, stat.ME) describe conditional tie weighting for win statistics, but the full manuscript text is SIKA-GP (arXiv 2605.26509, cs.LG) on sparse inducing kernels for GPs. None of the claimed identification argument, U-statistic theory, sandwich variances, simulations, or HF-ACTION reanalysis appears in the body. The central claim cannot be refereed from the abstract alone.
  2. Load-bearing identification (asserted only in the abstract): the claim that replacing the higher-priority genuine-tie indicator by its conditional probability given the observed pairwise data preserves the restricted-time win probabilities (rather than shifting the estimand) is the step on which the whole contribution rests. Without the derivation, censoring assumptions, and nuisance-model conditions, one cannot verify that fractional weights for censoring-induced ties leave the restricted-time estimand unchanged when lower-priority comparisons are informative.
  3. Uninspectable large-sample theory and inference: the abstract asserts two-sample U-statistics with estimated nuisance functions and sandwich variances for win ratio, net benefit, and win odds. The form of the kernel, the influence-function expansion, regularity conditions on the nuisance estimators, and the sandwich construction are not available in the supplied text, so asymptotic validity and variance correctness cannot be assessed.
  4. Uninspectable empirical support: efficiency gains under heavier censoring and longer restriction horizons, and the HF-ACTION death-first hospitalization reanalysis, are asserted but not present. Without design, data-generating mechanisms, competitor estimators, and numerical results, the practical claims cannot be checked.

Circularity Check

0 steps flagged

No circularity found: the restricted-time win estimand is pre-existing; conditional weights are proposed as an estimator of that same target, not as a redefinition of it.

full rationale

From the abstract (the only material that matches paper 2605.26507), the target is the same restricted-time win probabilities already used by existing all-or-nothing restricted win-statistic estimators. Conditional tie weighting replaces the unobserved higher-priority genuine-tie indicator by its conditional probability given the observed pairwise data so that partially observed pairs can contribute fractionally; the paper claims identification of those same restricted-time probabilities plus U-statistic asymptotics and sandwich variances. That is an identification/estimation claim, not a definitional loop: the estimand is not defined as whatever the weighted estimator produces, and no fitted parameter is renamed as a prediction of itself. Residual risk that misspecified nuisance models for the conditional genuine-tie probability could shift the estimand is a correctness/modeling concern, not circularity. The CACHEABLE full manuscript is a different paper (SIKA-GP, arXiv 2605.26509), so the win-statistic identification proof, U-statistic theory, and HF-ACTION design cannot be inspected here; that gap blocks verification of correctness but does not create a circular reduction by construction. No self-definitional step, fitted-input-as-prediction, load-bearing self-citation uniqueness claim, or renaming of a known result is quotable from the available text.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 1 invented entities

Abstract-only review of a biostatistics methods paper. Load-bearing content is the identification of restricted-time win probabilities under right censoring via conditional probabilities of higher-priority genuine ties, plus standard two-sample U-statistic asymptotics with estimated nuisances. Free parameters and invented entities cannot be enumerated from the abstract; the main modeling commitments are domain assumptions about hierarchical endpoints, independent censoring (likely), and correct nuisance estimation for the conditional tie probability.

free parameters (2)
  • restriction horizon / restricted-time window
    Restricted win statistics depend on a chosen time horizon; the abstract notes larger gains at longer horizons, so this design choice affects operating characteristics even if the estimand is defined relative to it.
  • nuisance models for conditional genuine-tie probability
    The method replaces an unobserved indicator by an estimated conditional probability; any parametric or semiparametric nuisance fit is a free modeling choice that the abstract does not specify.
axioms (4)
  • domain assumption Hierarchical composite comparison: patients are compared first on the highest-priority outcome and descend only on genuine ties.
    Core clinical-trial setup stated in the abstract; defines when lower-priority outcomes may contribute.
  • domain assumption Right censoring can induce apparent higher-priority ties even when a lower-priority comparison is already observed.
    Problem statement in the abstract; without this structure the proposed weighting is unnecessary.
  • ad hoc to paper The conditional probability of a higher-priority genuine tie given observed pairwise data identifies the contribution needed to preserve the restricted-time win probabilities.
    Central identification claim of the paper; asserted but not derived in the available abstract.
  • standard math Two-sample U-statistics with estimated nuisance functions admit standard large-sample expansions and sandwich variance estimators for win ratio, net benefit, and win odds.
    Abstract claims establishment of this theory; relies on classical U-statistic asymptotics under regularity conditions not stated here.
invented entities (1)
  • conditional tie weighting estimator no independent evidence
    purpose: Replace the unavailable higher-priority genuine-tie indicator by its conditional probability so partially observed pairs contribute fractionally to restricted win statistics.
    Primary methodological object introduced in the abstract; independent evidence would be the identification proof, simulations, and trial reanalysis, none of which are inspectable here.

pith-pipeline@v1.1.0-grok45 · 26390 in / 2974 out tokens · 28838 ms · 2026-07-12T15:56:52.836304+00:00 · methodology

0 comments
read the original abstract

Hierarchical composite endpoints are increasingly used in clinical trials to compare patients first on the most clinically important outcome and then, only when that comparison is tied, on lower priority outcomes. Under right censoring, a lower priority comparison may already be observed but still cannot contribute because the higher priority genuine tie required for descent through the hierarchy is not confirmed. Existing restricted win-statistic estimators address censoring by requiring such ties from higher priority to be observed as genuine ties. This all-or-nothing rule preserves the restricted-time estimand, but excludes pairs with censoring-induced ties even when their lower priority comparisons contain useful information. We propose conditional tie weighting, which replaces the unavailable higher priority genuine-tie indicator by its conditional probability given the observed pairwise data. The resulting estimator targets the same restricted-time win probabilities while allowing partially observed pairs to contribute fractionally when their lower priority comparison is informative. We establish identification and large-sample theory for the resulting two-sample U-statistics with estimated nuisance functions, and derive sandwich variance estimators for the win ratio, net benefit, and win odds. Simulations show substantial efficiency gains, especially under heavier censoring and longer restriction horizons. A reanalysis of the HF-ACTION trial illustrates how conditional tie weighting recovers information from censoring-induced ties in death-first hospitalization comparisons further apply our estimator to reanalyze a completed randomized clinical trial.

Figures

Figures reproduced from arXiv: 2605.26507 by Fan Li, Xi Fang.

Figure 1
Figure 1. Figure 1: Illustration of pairwise weighting rules under the IPCW estimator of Cui et al. (2025) and the proposed [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: An example of basis functions (L= 3). region [(m 1)2 l ,(m + 1)2l ] whenever l 1 . An example of lm (x) with L = 3 is illustrated in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Figure 2. Copula sensitivity of [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: SIKA-GP replaces dense GP inference with sparse inducing kernel approximations. A sparse feature mapping selects a small subset of activated basis functions (x) per input, which are processed by Bayesian feed-forward layers and aggregated additively, enabling efficient inducing kernel inference with near-linear complexity. and the effective random weights W are only D(L + 2)- dimensional, i.e., (L + 2)-dim… view at source ↗
Figure 3
Figure 3. Figure 3: Estimated net benefit (NB, top row), win ratio (WR, middle row), and win odds (WO, bottom row) as [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Parallelized lightweight forward process of SIKA-GP us￾ing TSI algorithm. “ ” denotes the inner product (i.e., convolution linear layer) on the last two dimensions [D; M]. The activated tensor indices in (X) are, therefore, obtained by offset t and concatenated with the first two indices (two globally activated 01, ψ02) J(X) = 2 6 4 1, 2, t+ (r/2 + 1) | {z } offset from to 3 7 5 . (16) Since only indices o… view at source ↗
Figure 5
Figure 5. Figure 5: CPU inference time of SIKA-GP. B: batch size; S: the number of MC samples; D: the dimension of features. (a) (B ; S; D)=(16;10;128). (b) (B ; S; D)=(16;10;512). (c) (B ; S; D)=(16;10;512). (d) (B ; S; D)=(128;10;512) [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: CUDA inference time of SIKA-GP. B: batch size; S: the number of MC samples; D: the dimension of features. During inference, we predict for a new x by similarly draw￾ing S samples from the learned variational distribution: E q(W ;b) [P(y jx , W , b)] 1 S SX s=1 P (y jx , W s , bs ) [PITH_FULL_IMAGE:figures/full_fig_p006_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: CUDA inference time of deep SIKA-GP [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 1 linked inside Pith

  1. [1]

    template basis function

    PMLR, 2013. Ding, L., Tuo, R., and Shahrampour, S. Generalization guarantees for sparse kernel approximation with entropic optimal features. InInternational Conference on Machine Learning, pp. 2545–2555. PMLR, 2020. Ding, L., Tuo, R., and Shahrampour, S. A sparse expansion for deep Gaussian processes.IISE Transactions, 56(5): 559–572, 2024. Duvenaud, D. K...

  2. [2]

    wavelet-like

    Let � � �������������� ������� . Suppose �� �,� � ��,� � is an orthonormal basis of � � . Then �� �� � ��� � �T�x� ����� ������ forms an orthonormal basis of��� ������������ �L �� ��������L �� � � � � �����L ���L ��. Proof. First, we prove Statement 1. Let ��� �� ����� and ���� � �� ���� ��. When ��� � and ��� �, this statement is ensured by the definitio...

  3. [3]

    calibration

    to implement GP modules. In classification tasks, we apply a softmax likelihood to normalize the output digits into probability distributions. DNNs are non-Bayesian models trained via negative log-likelihood loss, while DKL and DAK models are trained via ELBO loss. MNIST: The feature extractor is a simple CNN: ������ (1,32,3) ������� (32,64,3) ���������� ...