{"id":"e0aef839-86d3-4fd3-9809-0fdcdcb0777b","arxiv_id":"1908.07112","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MaxCombo, the maximum of four Fleming-Harrington weighted log-rank statistics, is proposed as a robust primary test and design basis for confirmatory trials with non-proportional hazards.","lead":"A cross-pharma group proposes using a combination of four weighted log-rank tests, called MaxCombo, as the primary analysis for confirmatory trials where treatment effects may change over time. The paper adds design, sample size, and reporting guidance, with worked examples from three cancer trials, giving practitioners a pre-specified robust alternative to the log-rank test.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Formal definition of MaxCombo is internally inconsistent: upper-tail max in Appendix A vs lower-tail cutoff -2.286 in Appendix D; as written the proposed test cannot be the robust test claimed.","rationale":"The reader's conditional verdict identified the simulated correlation matrix as the weakest assumption. That is a legitimate design-transparency concern, but it is not the most load-bearing issue: the null covariance of the four FH statistics can in principle be estimated from the observed data, and the paper's p-value formula in Appendix A does not force use of the design-stage matrix. The stronger problem is internal: the definition of the MaxCombo statistic and the p-value/boundary are mutually inconsistent. Appendix A is an upper-tail max; Appendix D uses a lower-tail cutoff. The examples are consistent with min-p combination, not max. This is exactly the kind of internal inconsistency that makes the central claim hard to evaluate. I agree with the reader that additional artifacts and transparency are needed, and the strong-null-2 (48.9% rejection) and PA3 close-miss results are acknowledged limitations that support a conditional rather than unconditional recommendation. Since the sign issue is likely a fixable notational error and the intended test is well known, I do not move the verdict to reject; I would keep the paper conditional on correcting the test-statistic definition and boundary and re-verifying the Table 2 and Section 3.3 results.","tokens_in":21810,"tokens_out":10896,"duration_ms":117086,"concrete_test":"Analytically: take the Appendix D null correlation matrix and compute q satisfying P(max(G0,0, G0,1, G1,0, G1,1) > q) = 0.025 for N4(0, Sigma). If q is positive (approx. +2.0) rather than -2.286, the Appendix D boundary is not the upper-tail MaxCombo boundary. Then simulate the Section 3.3 delayed-effect design exactly as written (negative G favors treatment, Zmax = max of the four G's, reject if Zmax < -2.286); the empirical power will be near zero, contradicting the claimed 90% power and the paper's own simulation summary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix A defines Zmax = max{G0,0, G0,1, G1,0, G1,1} and the one-sided p-value as P(Zmax > zmax) = 1 - Phi_4(zmax), which is an upper-tail test (positive G favors treatment). Appendix D instead solves Phi_4(Zcutoff) <= 0.025 and obtains Zcutoff = -2.286, a lower-tail critical value (negative G is significant). The two definitions cannot both be correct. If negative G favors treatment, the correct combination is the minimum (or the negative of the maximum), not the maximum; using 'max' would mask the most favorable component with the least favorable one and destroy power. If positive G favors treatment, the 2.5% upper-tail cutoff for the maximum of four correlated normals is about +2.0, not -2.286. The worked examples (Table 2) select G0,1 with p=0.002 and report MaxCombo p=0.005, consistent with a minimum-p/lower-tail procedure; this contradicts the 'max' definition. Because Section 3.1 uses the -2.286 boundary for sample size and the final p-value is tied to the same statistic, a practitioner following the equations literally cannot implement a test with the claimed operating characteristics. This is more fundamental than the correlation-matrix calibration issue: even with a perfectly known covariance, the test statistic and boundary are mutually inconsistent.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes the MaxCombo test, defined as the maximum of four correlated Fleming-Harrington weighted log-rank statistics (G0,0, G0,1, G1,1, G1,0), as a robust primary analysis for confirmatory clinical trials with potential non-proportional hazards (NPH). It presents an analysis workflow (test of the null, assessment of proportional hazards, and adaptive treatment-effect summaries), a design approach with sample-size calculation based on an adjusted significance boundary and an iterative simulation step, an interim-analysis strategy, and three illustrative applications from published oncology trials. The authors claim that MaxCombo controls type I error at 2.5% and provides robust power across PH, delayed effect, crossing survival, early separation, and mixtures of NPH patterns, and they frame the proposal as a 'straw man' guidance for discussion.","tokens_in":22181,"tokens_out":4751,"duration_ms":50654,"significance":"If the proposal were internally consistent and its operating characteristics were fully supported, the paper would provide valuable, practical guidance for an important methodological need: a pre-specified primary test with robust power under various non-proportional-hazards patterns while controlling type I error at the conventional level. The paper draws on established asymptotic results (Karrison 2016), provides worked examples and a concrete design algorithm, and explicitly acknowledges the limitations of single summary measures under NPH. However, the central test definition is marred by a sign/direction inconsistency between the formal definition and the numerical cutoffs and examples, and the main robust-power evidence is delegated to a companion paper by the same working group. These issues currently prevent the manuscript from fulfilling its stated goal of a directly implementable guidance.","major_comments":[{"comment":"The definition of the MaxCombo test is internally inconsistent. Appendix A defines Zmax = max{G0,0, G0,1, G1,0, G1,1} and gives the one-sided p-value as P(Zmax > zmax|H0) = 1 - Φ_4(zmax), which is an upper-tail procedure (large positive values are significant). Appendix D, however, solves Φ_4(Zcutoff, 0, ΣH0) ≤ 0.025 and obtains Zcutoff = -2.286, a lower-tail critical value for the maximum: rejection occurs when the maximum is below -2.286, i.e., when all four statistics are simultaneously strongly negative. For a maximum of four positively correlated normals, the one-sided 2.5% upper-tail cutoff is approximately +2.0, not -2.286. If negative values of the Fleming-Harrington statistics are meant to favor treatment, the appropriate combination is the minimum (or the negative of the maximum), not the maximum as written. The worked examples in Table 2 (e.g., IM211 row: selected G0,1 with p=0.002 and MaxCombo p=0.005) are consistent with a minimum-p/lower-tail procedure, contradicting the stated 'max' definition. Because Section 3.1 uses the -2.286 boundary for sample-size calculation and the final p-value is tied to the same statistic, a practitioner following the equations literally cannot implement a test with the claimed operating characteristics. This is a load-bearing inconsistency, independent of the correlation-matrix calibration issue.","section":"Appendix A vs. Appendix D"},{"comment":"The correlation matrix of the four Fleming-Harrington statistics is estimated from a single large simulated null trial with an assumed piecewise-exponential control survival, enrollment rate, and follow-up pattern, and is then treated as the true correlation matrix when computing the adjusted significance boundary and the final MaxCombo p-value. If the actual trial's null survival, censoring, or dropout pattern differs materially from the simulation assumptions, the boundary and p-value can be miscalibrated. The paper does not provide sensitivity analyses, an analytic characterization of the correlation as a function of the underlying censoring pattern, or a conservative fallback. Given that the proposal is for confirmatory regulatory use, the claim of type I error control at 2.5% needs stronger justification than a single simulation scenario.","section":"Section 3.1, Step 2 and Appendix D"},{"comment":"The core evidence for robust power and type I error control is delegated to Lin, Lin, Roychoudhury et al. (2020), a companion paper by the same working group. This manuscript itself reports only a limited set of strong-null simulations (Section 2.1.2) and the worked design example (Section 3.3). For a self-contained guidance paper that states the MaxCombo test 'fulfills the necessary regulatory standards,' the reader needs either a summary of the companion's operating characteristics, or at least a verification of the design example's power and type I error under the stated assumptions, rather than a citation. The single simulation check in Section 3.3 ('we confirm type I error and power using simulation') is not described in enough detail to be independently reproducible.","section":"Section 2.1.1 and Section 3.3"},{"comment":"The same direction inconsistency appears in the interim-analysis boundary equations. Appendix C states the final boundary condition as P(ZI > zI|H0) + P(ZI ≤ zI, ZF_max > zF_max|H0) ≤ 0.025, which is an upper-tail formulation with zI positive. Appendix D instead writes P(ZI < -2.34|H0) + P(ZI > -2.34, MF < zF|H0) ≤ 0.025, using a negative interim boundary and a lower-tail condition on MF. The sign of the interim eﬃcacy boundary and the direction of the final condition are thus inconsistent between the two appendices, which further complicates implementation.","section":"Appendix C and Appendix D"}],"minor_comments":[{"comment":"The phrase 'non-proportional hazard is a possibility' in the abstract and elsewhere should be 'non-proportional hazards' for consistency.","section":"Section 1.1"},{"comment":"The examples use reconstructed and unstratified data, while the published results are stratified; the paper acknowledges this, but it would be helpful to state explicitly that the reported p-values are not directly comparable to the original trial analyses.","section":"Section 2.3"},{"comment":"The sentence 'Further details of the sample size calculation including correlation matrix for null distribution are provided in Appendix D' is followed by another sentence about interim analysis; reordering for clarity would help.","section":"Section 3.3"},{"comment":"In the strong-null 1 description, 'The curves meet at the 36 month.' has a grammatical error and should probably read 'at 36 months.'","section":"Section 2.1.2"},{"comment":"The simultaneous confidence interval formula 'HRMaxCombo ± C*×SE(HRMaxCombo)' is written before defining C*; the notation should be introduced before the formula.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistency between the 'max' definition in Appendix A and the lower-tail cutoff in Appendix D is substantial and would require a careful rewrite of the test definition, the worked examples, and the design calculations. I would also encourage the editor to verify that the companion paper by Lin et al. (2020) uses a consistent definition, since the present manuscript relies heavily on it. The paper's framing as a 'straw man' does not reduce the need for internal consistency, especially in a methods paper aimed at practitioners."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper is the source of the widely used 'MaxCombo' proposal, but as written the formal definition does not hang together. Appendix A defines Zmax as the maximum of the four Fleming-Harrington statistics and gives the p-value as P(max > zmax), an upper-tail test where large positive values favor treatment. Appendix D then computes the design boundary by solving Phi_4(z) <= 0.025 and gets z = -2.286, a lower-tail critical value where large negative values reject. The two definitions cannot both be right. Table 2 tells you which one is intended: the selected component is G0,1 with p = 0.002 and the reported MaxCombo p is 0.005, which is the behavior of a minimum-p (or minimum-statistic) test, not a maximum. So the equations and the examples describe different tests. A practitioner following the paper literally cannot reproduce the claimed operating characteristics.\n\nWhat the paper does well: it lays out a practical weight set (0,0), (0,1), (1,0), (1,1), a sensible three-step reporting strategy for treatment-effect summaries after a robust test, and an iterative sample-size workflow with an interim-analysis boundary based on correlated FH statistics. It is also honest about the strong-null and late-crossing scenarios and offers a modified weight set that behaves better there. The examples are instructive, and the 'straw man' framing is appropriate.\n\nSoft spots: the operating-characteristic evidence for the main test is delegated to a companion paper; the strong-null-2 rejection rate of 48.9% and the PA3 PH near-miss cut against the 'suitable for confirmatory trials' claim. The sample-size calculation uses a simulated null correlation matrix as if it were known, so the calibration is only approximate if the real trial's event and censoring pattern differs. Those are addressable. The sign inconsistency is not minor and has to be fixed before anything else.\n\nWho it is for: trialist statisticians and regulators thinking about NPH primary analysis. Read it as a position paper, not a recipe. If I were the editor I would send it to peer review because it is influential and the flaws are fixable, but I would require the authors to reconcile the definition with the boundary and re-run the examples before publication.","headline":"MaxCombo is a genuinely useful straw-man proposal, but as written the test is internally inconsistent: the p-value is an upper-tail maximum while the design boundary is a lower-tail cutoff.","tokens_in":22727,"tokens_out":4613,"would_cite":false,"duration_ms":46929,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes the MaxCombo test—the maximum of four Fleming-Harrington weighted log-rank statistics—as a primary analysis for confirmatory trials with non-proportional hazards, with design and sample-size guidance to match.","keywords":["non-proportional hazards","MaxCombo test","Fleming-Harrington weighted log-rank test","combination test","confirmatory trial design","sample size calculation","interim analysis","type I error control"],"falsifier":"Simulate a null trial with no treatment effect but with a survival or censoring pattern that departs from the assumed piecewise-exponential setup (for example, a cure fraction or heavy late censoring), apply the paper's Appendix D boundary and MaxCombo p-value calculation, and check whether the one-sided rejection rate stays at 2.5%; a material excess would show the calibration depends on the assumed null correlation matrix.","tokens_in":21604,"feed_emoji":"📊","tokens_out":14052,"duration_ms":117440,"temperature":0.7,"pith_summary":"The paper argues that a single log-rank test and a hazard ratio can fail badly when treatment effects emerge late, cross over time, or fade, and that in such settings no one number adequately describes the treatment effect. It proposes MaxCombo, the maximum of four Fleming-Harrington weighted log-rank statistics, as a primary analysis test with robust power across proportional and non-proportional hazards. The authors give a full design-and-analysis package: how to compute the p-value from the joint normal distribution, how to adjust the significance level for the four correlated tests, how to size a trial by simulation, and how to set interim and final boundaries. Three reconstructed oncology trials illustrate that MaxCombo can catch a crossing-survival benefit that the log-rank test misses, while losing only a little power when proportional hazards hold.","feed_headline":"MaxCombo test keeps power when survival curves cross or delay","feed_subtitle":"The maximum of four weighted log-rank tests protects the 2.5% error bound while catching late and crossing effects.","key_machinery":"The central object is the MaxCombo statistic $Z_{max} = \\max\\{G^{0,0}, G^{0,1}, G^{1,1}, G^{1,0}\\}$, where each $G^{\\rho,\\gamma}$ is a Fleming-Harrington weighted log-rank statistic; the weights emphasize early events ($\\rho$), late events ($\\gamma$), or both. The load-bearing identity is the covariance formula $\\eta_{ij} = V(G_{(\\rho_i+\\rho_j)/2,(\\gamma_i+\\gamma_j)/2}) / \\sqrt{V(G_{\\rho_i,\\gamma_i}) V(G_{\\rho_j,\\gamma_j})}$, which turns the joint null distribution of the four statistics into an estimable multivariate normal and lets the design compute an adjusted boundary without per-trial simulation. Around this identity the paper builds an iterative design procedure: simulate one large null trial to estimate the correlation matrix, solve for the per-component significance level, size each of the four tests with a published weighted-log-rank sample-size formula, and confirm the operating characteristics by simulation. For trials with an interim analysis, the same framework uses the independent-increments property to correlate the interim log-rank statistic with the final MaxCombo statistic.","core_discovery":"The paper's central claim is that the MaxCombo test—the maximum of four correlated Fleming-Harrington weighted log-rank statistics $G^{0,0}$, $G^{0,1}$, $G^{1,1}$, and $G^{1,0}$—is suitable as the primary analysis test in a confirmatory trial when non-proportional hazards are a real possibility. Under the null hypothesis the four statistics are asymptotically multivariate normal, so a one-sided p-value can be obtained by integrating a four-dimensional normal density above the observed maximum, and the significance boundary can be adjusted using their correlation matrix rather than a conservative Bonferroni correction. Simulation and three reconstructed trial examples are used to argue that the test controls type I error at 2.5%, has strong power for delayed, crossing, early-separation, and mixed non-proportional-hazard patterns, and loses only modest power relative to the log-rank test under proportional hazards. The paper pairs the test with a three-step analysis strategy (test the null, assess proportional hazards, then report either the hazard ratio or a set of time-dependent summaries) and a simulation-based design procedure for sample size and interim boundaries.","pith_inferences":["One implication not developed in the paper is that the practical bottleneck for MaxCombo will be regulator agreement on the simulation plan and the null correlation matrix in advance, not the test statistic itself.","A natural extension would be to let the data at the final analysis determine the correlation matrix from the observed event and censoring pattern instead of fixing it at the design stage, and to study how sensitive the reported p-value is to that choice.","The same maximum-of-correlated-weighted-log-rank construction could be applied to a larger or differently chosen set of weights; everything needed for the p-value would follow from the same covariance identity, and the operating characteristics would need rerunning.","A testable refinement motivated by the paper's examples is to treat the MaxCombo test as a gatekeeper and then quantify how much the treatment-effect estimate changes across the four component weights, exposing whether the smallest p-value also corresponds to a clinically meaningful effect."],"forward_implications":["A trial facing uncertain non-proportional hazards can pre-specify MaxCombo as the primary test and still control one-sided type I error at 2.5%.","In the delayed-effect design example, MaxCombo needs 472 patients and 372 events versus 690 patients and 544 events for the log-rank test, a substantial saving.","Designs should specify minimum follow-up (roughly twice the control median) alongside event count, because MaxCombo power depends on follow-up, not just events.","For a trial with an interim log-rank analysis, the final MaxCombo boundary can be adjusted using the independent-increments correlation between the interim statistic and the four final statistics.","Under proportional hazards with a modest benefit, MaxCombo can lose power relative to the log-rank test, so the paper recommends reserving it for settings where non-proportional hazards are plausible."],"supporting_citations":[{"why":"Supplies the asymptotic multivariate normal joint distribution of weighted log-rank statistics and the correlation formula underlying MaxCombo p-values and boundaries.","marker":"Karrison (2016)"},{"why":"Introduces the combination of weighted log-rank statistics that motivates the MaxCombo maximum.","marker":"Lee (2007)"},{"why":"Provides the comparative simulation study that supplies the type I error and robust-power operating characteristics claimed for MaxCombo.","marker":"Lin, Lin, Roychoudhury et al. (2020)"},{"why":"Gives the sample-size formula for weighted log-rank tests with Fleming-Harrington weights used in the design's initial sample-size calculation.","marker":"Hasegawa (2014)"},{"why":"Defines the modestly weighted log-rank test and the strong-null 1 scenario used to stress-test MaxCombo.","marker":"Magirr and Burman (2019)"},{"why":"Provides the strong-null 2 scenario and the critique of weighted log-rank tests that the paper addresses with a modified MaxCombo.","marker":"Freidlin and Korn (2019)"},{"why":"Supplies the independent-increments property used to correlate the interim log-rank statistic with the final MaxCombo statistic.","marker":"Tsiatis (1982)"},{"why":"Provides the alpha-spending function used to set the interim efficacy boundary.","marker":"Demets and Lan (1994)"},{"why":"Defines the $G^{\\rho,\\gamma}$ class of weighted log-rank statistics of which the four MaxCombo components are members.","marker":"Harrington and Fleming (1982)"}],"fun_headline_variants":["MaxCombo tames crossing and delayed survival curves","Combination test keeps power for non-proportional hazards","MaxCombo test preserves power when survival curves cross","Robust test for trials with non-proportional hazards","MaxCombo test controls error for crossed survival data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design treats the correlation matrix of the four test statistics, estimated from one large simulated null trial under assumed piecewise-exponential survival, enrollment, and follow-up, as the true correlation matrix of the actual trial; if the real null survival or censoring pattern differs materially, the significance boundary and final p-value can be miscalibrated.","fun_headline_variants_meta":{"raw":{"variants":["MaxCombo tames crossing and delayed survival curves","Combination test keeps power for non-proportional hazards","MaxCombo test preserves power when survival curves cross","Robust test for trials with non-proportional hazards","MaxCombo test controls error for crossed survival data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000912,"raw_usage":{"total_tokens":3953,"prompt_tokens":1015,"completion_tokens":2938,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":2862}},"tokens_in":631,"tokens_out":2938,"duration_ms":25243,"temperature":1.0,"reasoning_tokens":2862,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:25:52.099433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a null trial with no treatment effect but with a survival or censoring pattern that departs from the assumed piecewise-exponential setup (for example, a cure fraction or heavy late censoring), apply the paper's Appendix D boundary and MaxCombo p-value calculation, and check whether the one-sided rejection rate stays at 2.5%; a material excess would show the calibration depends on the assumed null correlation matrix.","supporting_citations":[{"cited_title":"Versatile tests for comparing survival curves based on weighted log-rank statistics","cited_arxiv_id":null,"evidence_quote":"Supplies the asymptotic multivariate normal joint distribution of weighted log-rank statistics and the correlation formula underlying MaxCombo p-values and boundaries."},{"cited_title":"On the versatility of the combination of the weighted log-rank statistics","cited_arxiv_id":null,"evidence_quote":"Introduces the combination of weighted log-rank statistics that motivates the MaxCombo maximum."},{"cited_title":"Alternative analysis methods for time to event endpoints under nonproportional hazards: A comparative analysis","cited_arxiv_id":null,"evidence_quote":"Provides the comparative simulation study that supplies the type I error and robust-power operating characteristics claimed for MaxCombo."},{"cited_title":"Sample size determination for the weighted log-rank test with the fleming–harrington class of weights in cancer vaccine studies","cited_arxiv_id":null,"evidence_quote":"Gives the sample-size formula for weighted log-rank tests with Fleming-Harrington weights used in the design's initial sample-size calculation."},{"cited_title":"Modestly weighted logrank tests","cited_arxiv_id":null,"evidence_quote":"Defines the modestly weighted log-rank test and the strong-null 1 scenario used to stress-test MaxCombo."},{"cited_title":"Methods for accommodating nonproportional hazards in clinical trials: Ready for the primary analysis? Journal of Clinical Oncology 2019; 37(35): 3455--3459","cited_arxiv_id":null,"evidence_quote":"Provides the strong-null 2 scenario and the critique of weighted log-rank tests that the paper addresses with a modified MaxCombo."},{"cited_title":"Repeated significance testing for a general class of statistics used in censored survival analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the independent-increments property used to correlate the interim log-rank statistic with the final MaxCombo statistic."},{"cited_title":"Interim analysis: The alpha spending function approach","cited_arxiv_id":null,"evidence_quote":"Provides the alpha-spending function used to set the interim efficacy boundary."},{"cited_title":"A class of rank test procedures for censored survival data","cited_arxiv_id":null,"evidence_quote":"Defines the $G^{\\rho,\\gamma}$ class of weighted log-rank statistics of which the four MaxCombo components are members."}],"review_version":1}