Pith. sign in

REVIEW 4 major objections 4 minor 5 references

Advancing Waterfall Plots for Cancer Treatment Response Assessment through Adjustment of Incomplete Follow-Up Time

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A weighted Kaplan-Meier adjustment projects ongoing-study waterfall plots to their expected final form, so interim efficacy can be read against completed historical controls.

desk verdict A genuinely useful adjustment to interim waterfall plots with a clean derivation, but the validation leans on an ad hoc filter and an untested independence assumption. read the letter →

arxiv 2506.07365 v1 pith:H5CCVDJQ submitted 2025-06-09 stat.AP

classification stat.AP MSC 62N0162P1062F15
keywords waterfallplotsearlyphaseoncologytumorresponseassessmentweightedKaplan-Meierincompletemultinomialmodelcensoringhistoricalcontrolcomparisoninterimanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to fix a fairness problem in early-phase oncology: an ongoing trial's waterfall plot shows each patient's best tumor shrinkage so far, which for patients still on treatment understates the shrinkage they may eventually reach. The authors argue that a waterfall plot is a rotated survival curve of best tumor-size change, so they adjust an interim Kaplan-Meier curve by weighting each ongoing patient by the probability that their current best response is already their final best response. That probability is estimated through an incomplete multinomial model for the scan at which the final best response will occur. In a Phase 2 data cut and across 300 data replications, the adjusted waterfall curve sits closer to the true final-best-response curve than the unadjusted curve. If the method holds, interim experimental arms can be compared with historical controls from completed studies without waiting for full follow-up or needing the controls' individual patient data.

What carries the argument

The central object is the weighted Kaplan-Meier curve of the negative final best tumor size change, $S_w(z) = \prod_{u \le z} \left(1 - \frac{dN(u)\,p_i(u)}{R(u)}\right)$, where each event's contribution is multiplied by the patient's probability that the current best response is final. Those per-patient probabilities come from an incomplete multinomial model for the scan index $U_i$ of the final best tumor size change: conditional on interim data, the final best scan is drawn from the possible future scans $S_i$ with category probabilities $\theta_k$, so $p_i = \theta_{U_i} / \sum_{k\in S_i} \theta_k$, and $\theta$ is estimated by a Bayesian Gibbs sampler with a Dirichlet prior. Rotating and reversing this weighted survival curve reproduces the waterfall plot, and the paper shows the weighted curve equals a probability-weighted average of conventional KM curves over all possible event/censoring configurations of ongoing patients.

What would settle it

Using a completed trial where final best responses are known, cut the data at an early enrollment fraction, apply the adjustment, and compare the adjusted interim waterfall to the actual final waterfall across many random cuts; if, after conditioning on current response depth, the estimated probability of further improvement still differs systematically by depth, the common-distribution assumption fails and the adjusted curve will not center on the ground truth.

Watch

Extended reading notes

Core claim

The paper's central claim is that a waterfall plot from a study with incomplete follow-up can be projected to approximate the waterfall plot that would be seen with sufficient follow-up by treating 'the final best tumor size change equals the current best tumor size change' as an event and 'further improvement is still possible' as censoring. Under this framing, each ongoing patient enters the interim Kaplan-Meier curve for negative best tumor size change with a weight equal to the estimated probability that the current scan is the final best scan, $p_i = \theta_{U_i} / \sum_{k \in S_i} \theta_k$, with $\theta$ estimated by an incomplete multinomial model for the scan index of the final best response. Rotating and reversing this weighted survival curve yields the adjusted waterfall plot. The paper demonstrates on a Phase 2 trial that the adjusted curve is closer to the true fBTSC waterfall, while the unadjusted curve systematically underestimates efficacy; across 300 replications, adjusted curves center around the ground truth. The intended consequence is that adjusted waterfall plots from ongoing trials are suitable for comparison with historical controls from completed studies, using only scan-index information and no individual-level control data.

Load-bearing premise

The load-bearing assumption is that the scan at which a patient will reach their final best tumor response follows one common distribution shared by all patients, independent of how deep their current best response already is; if patients with deeper early responses improve on a different timetable, the estimated improvement probabilities are off and the adjusted plot can still miss the true final waterfall.

Editorial extensions

If this is right

  • Interim waterfall plots from ongoing experimental arms can be compared directly with historical controls from completed studies, with no patient-level control data required.
  • The adjusted curve approximates the final best-response waterfall, reducing the systematic underestimation of tumor response in ongoing patients.
  • Pointwise confidence intervals from the weighted Kaplan-Meier fit can be rotated into uncertainty bands for the adjusted waterfall plot, giving a sense of how reliable the projection is.
  • The method needs only scan-index information, not a model for individual future tumor trajectories, making it practical for early-phase settings.
  • The weighted curve interpretation means the adjustment is equivalent to averaging over all possible outcomes for ongoing patients, which clarifies why it should center on the final distribution if the assumptions hold.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The paper's recommended filter, which excludes ongoing patients whose last two scans are unchanged or who have had any tumor increase, is an implicit acknowledgment that the shared-distribution model can over-predict improvement; a formal calibration study of that filter would be a useful next step.
  • Inference: The same weighted-survival rotation could be applied to any 'best value so far' endpoint where the final best is censored, such as depth of response in other disease settings or nadir-type biomarkers, not just tumor size change.
  • Inference: Because the model uses scan-index categories, applying it across studies with different scan schedules requires aligning the category sets $S_i$; otherwise the estimated event probabilities are not directly comparable.
  • Inference: A stronger validation would be to run the adjustment on several completed trials, artificially truncate follow-up at multiple enrollment fractions, and check that the adjusted curves remain centered on the true final waterfall across all fractions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a method to adjust interim waterfall plots in oncology trials so that they approximate how the plot would look with sufficient follow-up. The approach frames the final best tumor size change (fBTSC) relative to the current best tumor size change (cBTSC) as an 'event'/'censoring' problem, estimates the probability of no further improvement through an incomplete multinomial model for the scan index of fBTSC, and then builds a weighted Kaplan-Meier curve of the negative fBTSC, which is rotated to obtain an adjusted waterfall plot. The method is applied to a Phase 2 study, with a replication exercise, and compared against the completed-study ground truth.

Significance. If the proposed adjustment is valid, it would allow interim waterfall plots to be compared with historical controls from completed studies without requiring individual-level control data, which is an important practical need in early-phase oncology decision-making. The scenario-averaging interpretation of the weighted Kaplan-Meier curve is a useful conceptual bridge between survival analysis and waterfall plots, and the paper makes a concrete algorithmic proposal with a real-data illustration. However, the validity of the method rests on a strong and unvalidated assumption: that the probability of future tumor-response improvement does not depend on the depth of the current best response or the trajectory leading to it. The real-data demonstration also relies on an ad hoc filter that is introduced after the results are shown, which weakens the evidential value of the reported performance. The replication study shuffles start dates but preserves each patient's BTSC trajectory, so it cannot detect violations of the key assumption.

major comments (4)
  1. [Section 2.2, Eq. (1)-(2)] The incomplete multinomial model assumes a single common category distribution θ for the fBTSC scan index, independent of the magnitude of the current best tumor size change (Z_i) and other patient characteristics. The event probability p_i is computed as θ_{U_i} / Σ_{k∈S_i} θ_k, which depends only on the scan of cBTSC and the allowable future scan set, not on how deep the current response is. If patients with deeper responses have less room for further improvement, these probabilities are misspecified and the adjusted waterfall curve will be biased. This is a load-bearing assumption because the entire adjustment in Eq. (4) is built on p_i as a known or correctly estimated event probability. The paper does not provide evidence that this assumption holds, and the replication analysis in Section 4.1 shuffles treatment start dates while holding each patient's BTSC trajectory (and hence the cBTSC-depth/future-improvement relationship) fixed, so it cannot detect this misspecification.
  2. [Section 4.2 (Computational Considerations)] The 'filter' for ongoing patients is a post-hoc addition that is not prespecified and is not derived from the statistical model. The filter forces p_i=1 for patients whose most recent two scans are unchanged or who have experienced any tumor increase, effectively overriding the multinomial-model probabilities for a substantial subset of patients. Figures 3 and 4 report results based on this filtered version, so the demonstrated performance is not that of the method described in Sections 2-3. The paper should present unfiltered results and clearly pre-specify any filter in the Methods section, or better, incorporate the trajectory information into the model itself (e.g., by making θ depend on covariates such as current response depth and recent scan changes). As it stands, the real-data validation is not a clean test of the proposed methodology.
  3. [Section 4.1 (Analysis Results)] The agreement between the estimated θ and the observed marginal frequencies of fBTSC scan times (e.g., estimated 0.35 vs. actual 0.32 for scan 1) is taken as evidence that the model is performing well. However, this marginal agreement is not sufficient for the method's validity. The adjusted waterfall curve depends on the patient-specific event probabilities p_i, which are conditional quantities. A model could match the marginal distribution of fBTSC scan times while producing systematically incorrect p_i's for patients with different cBTSC depths. The paper should include a calibration check, for example, grouping patients by predicted p_i and comparing the observed proportion of 'events' (no further improvement) to the average predicted probability. Such a diagnostic would directly test the model's conditional specification.
  4. [Section 3 (Eq. 4) and Section 4.1 (Figure 3)] The confidence intervals in Figure 3 are obtained by applying the waterfall transformation to the survfit output for the weighted KM curve, treating the fitted weights (derived from the estimated θ and the filter) as fixed. This ignores the uncertainty in estimating θ and in applying the ad hoc filter, so the reported pointwise CIs are likely too narrow. The paper should either implement a bootstrap that accounts for the full estimation procedure, or explicitly state that the displayed intervals are conditional on the estimated weights. This is material for practical interpretation, because the figure claims that the adjusted curve's CI 'effectively covers' the ground truth.
minor comments (4)
  1. [Section 2.2] The text states 'fBTSC is less than or equal to the cBTSC (i.e., Z_i ≥ Z_i)' but the intended comparison is between the negative fBTSC (denoted Z̃_i) and the negative cBTSC (denoted Z_i); the tildes are missing. As written, the inequality is an identity, which is not informative.
  2. [Section 3 (R code example)] The R code example uses Z values that are inconsistent with the definition of Z_i as the negative of the tumor size change. In the example, a patient with a 30% reduction is listed as Z=-30 in the data frame, but under the paper's definition it should be Z=30. This inconsistency will confuse readers trying to reproduce the method.
  3. [Section 4.2 and throughout] There are several typographical errors and informal phrases that should be corrected, e.g., 'patietns' for 'patients', 'O btaining' for 'Obtaining', and the broken line break in the likelihood equation. The paper would also benefit from a clear statement that the filter in Section 4.2 was selected based on the authors' experience with 'various real datasets', which currently suggests a lack of pre-specification.
  4. [Section 1 and Section 4.1] The paper claims that the adjusted plots are 'suitable for comparison with historical controls', but the validation only compares against the same study's completed fBTSC, not against an external historical control. While the method may help account for follow-up differences, it does not address other between-trial differences (e.g., patient population, scan schedules). The claim should be softened or demonstrated with a case study involving actual historical controls.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the adjustment is derived from an incomplete multinomial model fit to interim data and validated against external fBTSC ground truth.

full rationale

The paper's central derivation proceeds as follows: interim observables (the scan index U_i at which cBTSC occurred and the feasible final-scan set S_i) enter an incomplete multinomial likelihood (Eq. 3); a Gibbs sampler estimates category probabilities theta; Eqs. (1)-(2) convert theta into per-patient event/censoring probabilities; Eq. (4) turns those probabilities into weights in a Kaplan-Meier curve; rotation/reversal yields the adjusted waterfall plot. The target fBTSC distribution is never used to set theta or p_i in the method as derived; it appears only as an external benchmark in Section 4.1 (Figure 3) and in the 300-replication check, where the actual fBTSC curves are the object being predicted, not an input to the estimator. The waterfall-survival relationship credited to Sun (2021) is self-cited but is also demonstrated in the present paper via Figure 1 and the toy example, so it is independent, mathematical support rather than a load-bearing self-citation. Section 4.2's post-hoc filter and Section 2.2's independence assumption (p_i depends on scan index, not on response depth) are modeling limitations or robustness heuristics; they create potential bias but do not make the claimed prediction equivalent to its inputs by construction. No equation or fitted constant is renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work. The honest finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the multinomial model for the timing of final best response, plus the waterfall-survival relationship. The filter is an additional ad hoc rule. No new physical or conceptual entities are introduced beyond pseudo patients, which are a computational device.

free parameters (3)
  • theta_k (category probabilities) = θ = (0.35, 0.2, 0.25, 0.1, 0.1) in real data
    Probabilities that fBTSC occurs at each scan; estimated from interim data via incomplete multinomial model. They drive the event probabilities p_i in equations (1)-(2).
  • K (number of categories) = K = J+1 where J = max scan of cBTSC among ongoing
    Modeling choice to merge scans after J+1; affects the multinomial categories but not the p_i calculations due to summation.
  • Dirichlet prior α = α = (1,...,1)
    Flat prior chosen for the Bayesian estimator; a modeling choice.
assumptions (4)
  • standard math A waterfall plot of BTSC is a rotated and reversed survival function of -BTSC (Sun et al. 2021).
    Used in Section 3 to justify transforming weighted KM curves into waterfall plots.
  • domain assumption The scan index of fBTSC for ongoing patients follows a multinomial distribution with category probabilities proportional to a common θ over the possible scan set S_i.
    Introduced in Section 2.2, equations (1)-(2). This assumes the timing of future best response is independent of the current response magnitude.
  • ad hoc to paper Ongoing patients who pass the filter (not two stable scans and no tumor increase) have a nonzero chance of further tumor reduction; those who fail are forced to p_i=1.
    Section 4.2 introduces this rule based on practical experience, not derived from the model.
  • domain assumption Censoring (future improvement) is independent of the observed cBTSC value conditional on S_i.
    Underlies the weighted KM estimator; not tested in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advancing Waterfall Plots for Cancer Treatment Response Assessment through Adjustment of Incomplete Follow-Up Time." pith.science (2026). https://pith.science/paper/H5CCVDJQ

@misc{pith2026250607365,
  author       = {Pith},
  title        = {Pith review of: Advancing Waterfall Plots for Cancer Treatment Response Assessment through Adjustment of Incomplete Follow-Up Time},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H5CCVDJQ}},
  note         = {Machine review of arXiv:2506.07365}
}
read the original abstract

Waterfall plots are a key tool in early phase oncology clinical studies for visualizing individual patients' tumor size changes and provide efficacy assessment. However, comparing waterfall plots from ongoing studies with limited follow-up to those from completed studies with long follow-up is challenging due to underestimation of tumor response in ongoing patients. To address this, we propose a novel adjustment method that projects the waterfall plot of an ongoing study to approximate its appearance with sufficient follow-up. Recognizing that waterfall plots are simply rotated survival functions of best tumor size reduction from the baseline (in percentage), we frame the problem in a survival analysis context and adjust weight of each ongoing patients in an interim look Kaplan-Meier curve by leveraging the probability of potential tumor response improvement (i.e., "censoring"). The probability of improvement is quantified through an incomplete multinomial model to estimate the best tumor size change occurrence at each scan time. The adjusted waterfall plots of experimental treatments from ongoing studies are suitable for comparison with historical controls from completed studies, without requiring individual-level data of those controls. A real-data example demonstrates the utility of this method for robust efficacy evaluations.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

5 extracted references · 5 canonical work pages

  1. [1]

    waterfall

    Introduction In oncology drug development, waterfall plots have emerged as a widely used tool to visualize individual patients' tumor size measurements. These graphical summaries, as demonstrated by early adopters (e.g., Campbell et al., 2007, Socinski et al., 2008, and Kwak et al., 2010), provide an intuitive and reliable method for assessing the efficac...

  2. [2]

    event” and “censoring

    Probabilities of “event” and “censoring” In this section, we set up the problem and establish a statistical model to estimate the probabilities of "event" and “censoring". 2.1 Setup Let us start with variables which are observable from the interim data. For an individual patient i, let 𝑍𝑍𝑖𝑖 denote the -cBTSC, and 𝑈𝑈𝑖𝑖 denote at which post-treatment scan (...

  3. [3]

    event" and Confidential

    Weighted Kaplan-Meier curves Understanding the relationship between waterfall plots and survival functions is crucial for our proposed adjustments. As demonstrated in Sun’s paper (2021), let 𝑋𝑋 be a generic random variable; then the waterfall plots of 𝑋𝑋 is essentially the survival function of −𝑋𝑋 after rotation and reverse. To illustrate, consider Plot (...

  4. [4]

    ground truth

    Real Data Analysis In this section, we apply the proposed method to a Phase 2 oncology clinical study to retrospectively assess early efficacy signal and facilitate the comparison against historical controls. By projecting waterfall plot as if follow-up were sufficient, this method could have helped evaluate whether the experimental treatment demonstrated...

  5. [5]

    apple-to-apple

    Discussion The conventional waterfall plots of ongoing studies with limited follow-up, while insightful, cannot take into account the potential for improved tumor response in ongoing patients. Our proposed method addresses this gap by providing a direct adjustment to the waterfall plot, projecting how the plot might look with sufficient follow-up. Our app...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.