{"id":"3ffaa388-5efd-4f86-b9e5-3ab2b213bd39","arxiv_id":"2506.03086","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A combination-trial design method is proposed, but the power and allocation calculations rest on an inverted variance formula.","lead":"This paper proposes a statistical framework for designing early-phase platform trials that test combination therapies, including a correlation-aware multiple-testing procedure and sample-size calculations. The authors also derive allocation rules intended to maximize power, but a central algebraic error in the noncentrality formula invalidates the claimed optimal allocations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (7) inverts the variance denominator, so the optimal-allocation and sample-size results rest on a noncentrality parameter that is the reciprocal of the correct one; this invalidates the paper's central design claim.","rationale":"The reader's rationale already identifies the reciprocal error, and an independent re-derivation of Eq. (7) from the definitions in Section 2.3 and Appendix 2 confirms it. This is the most load-bearing issue: every allocation-optimization result in Sections 2.3–2.5 and the real-data design recommendations depend on W_AB and W_B as written. The FWER control via critical values from Eq. (2) is a separate piece and is less affected, so the rejection should be understood as 'central contribution invalid as stated' rather than 'no salvageable components.' I also considered the reader's weakest assumption: nonzero arm-level correlations in a randomized parallel-arm trial are not justified because independent patient groups give zero correlation between sample means; if those correlations are zero, Eq. (2) reduces to the classical Dunnett correlation. This is a real second concern, but the reciprocal error is sufficient to reject and is the better target for a decisive test. The proposed check—re-solving the max-min problem with the correct reciprocal forms—would settle whether Eq. (12) is a valid optimum; it is not. Therefore the reader's REJECT verdict stands unchanged.","tokens_in":26085,"tokens_out":9151,"duration_ms":99043,"concrete_test":"Set s = 2, ρ_AB,A = ρ_AB,B = 0, and re-solve max_p min{4/(1/p_AB + 1/p_A), 1/(1/p_A + 1/p_B)} subject to p_A + p_B + p_AB = 1. Compare the maximizing allocation with Eq. (12). If the correct optimum differs from Eq. (12), or if Eq. (12) fails to equalize the correctly defined W_AB and W_B, the published closed-form allocation is disproved. A supplementary Monte Carlo power comparison at the two allocations would confirm which allocation dominates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 defines the Wald noncentrality parameters as W_AB = δ_AB² / Var(Ȳ_AB − Ȳ_A) and W_B = δ_B² / Var(Ȳ_B − Ȳ_A). From the model in the same section, Var(Ȳ_AB − Ȳ_A) = (σ²/N)[1/p_AB + 1/p_A − 2ρ_AB,A/√(p_AB p_A)] and Var(Ȳ_B − Ȳ_A) = (σ²/N)[1/p_A + 1/p_B]. Therefore the correct forms are W_AB = (N s²δ²/σ²) / [1/p_AB + 1/p_A − 2ρ_AB,A/√(p_AB p_A)] and W_B = (N δ²/σ²) / [1/p_A + 1/p_B]. Equation (7) multiplies by the variance bracket instead of dividing by it, and Appendix 2 repeats this inversion when it claims the result matches Eq. (7). Because the max-min problem (8), the equating step (10), the closed-form allocation (12), and the numerical optimization in Appendix 3 all use these W expressions, the claimed optimal allocations and the resulting sample-size recommendations do not follow from the stated noncentrality parameters. The error is not cosmetic: with s ≠ 1, the allocations from Eq. (12) do not equalize the correctly defined W_AB and W_B, so they solve a different optimization problem and can reduce power relative to the true max-min solution. The FWER-controlling critical-value computation based on Eq. (2) is less affected, but the paper's main design contribution is invalid as stated. A separate, independent concern is that nonzero arm-level correlations ρ_AB,A and ρ_AB,B are assumed for a randomized parallel-arm trial; if the arms enroll disjoint patient groups, these correlations are zero by design and Eq. (2) collapses to the classical Dunnett correlation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a statistical framework for designing early-phase platform trials of combination therapies. The main contributions are a generalized Dunnett's procedure that incorporates correlations between the combination arm and its monotherapy/control arms, an allocation-ratio optimization driven by a synergy parameter, sample-size determination via Monte Carlo binary search, an extension to K substudies, simulation studies, and a real-data application using patient-derived xenograft (PDX) data. An open-source R package is provided.","tokens_in":26544,"tokens_out":12874,"duration_ms":117744,"significance":"If the proposed methods were correct, they would offer a practical toolkit for combination-trial design, combining multiplicity control with synergy-informed sample allocation and translational data integration. The manuscript is clearly organized, contains extensive simulations, and ships an R package with documentation. However, the central power/allocation results contain algebraic errors that invalidate the main methodological claims, and the assumed inter-arm correlations are not justified for a randomized parallel-arm design. These issues are load-bearing, so the contribution as stated cannot be accepted.","major_comments":[{"comment":"The Wald noncentrality parameters are defined as W = δ² / Var(mean difference). From the variance expressions in Appendix 2, Var(Ȳ_AB − Ȳ_A) = (σ²/N)[1/p_AB + 1/p_A − 2ρ_AB,A/√(p_AB p_A)] and Var(Ȳ_B − Ȳ_A) = (σ²/N)(1/p_A + 1/p_B). Therefore W_AB must equal (N s² δ²/σ²) divided by the first bracket, and W_B must equal (N δ²/σ²) divided by the second bracket. Equation (7) and the final display of Appendix 2 multiply by these brackets, which inverts the noncentrality parameters. This error propagates into the max-min program (8), the equality condition (10), the closed-form allocation (12), the numerical optimization in Appendix 3, the simulations in Section 3.3, and the sample sizes in Table 2. A concrete check: for s = 2 and ρ_AB,A = 0, the allocation from Eq. (12) gives W_AB* = 4(1/0.211 + 1/0.366) ≈ 29.9 and W_B* = 1/0.366 + 1/0.423 ≈ 5.1, so the two noncentrality parameters are not equal and the claimed max-min solution does not equalize them.","section":"§2.3, Eq. (7); Appendix 2"},{"comment":"The derivation and the proposed generalized Dunnett procedure assume nonzero endpoint correlations ρ_AB,A and ρ_AB,B between the combination arm and the other arms. In a randomized parallel-arm platform trial, patients are randomized to disjoint arms, so the sample means from different arms are independent; consequently ρ_AB,A = ρ_AB,B = 0 by design, and Eq. (2) reduces to the classical Dunnett correlation. If the intended setting is instead a paired design in which the same experimental unit receives multiple treatments (as in the PDX data), the manuscript does not describe how such a trial would be randomized, how the analysis would account for the pairing, or how the correlations would be estimable in a clinical trial. This assumption is load-bearing: without nonzero inter-arm correlations, the claimed generalization of Dunnett's procedure disappears.","section":"§2.1, Eq. (2); Appendix 1"},{"comment":"The derivation of the closed-form allocation contains an algebraic error. Substituting p_B = 1 − x − y into the equality s²(1/x + 1/y) = 1/y + 1/(1−x−y) gives s²(x+y)(1−x−y) = x(1−x), which expands to y² + (2x−1)y + (1−1/s²)x² + (1/s²−1)x = 0 (after dividing by s²), not y² + (2x−1)y + (1−s²)x² + (s²−1)x = 0 as claimed. The roots y = 1−(s+1)x and y = (s−1)x therefore do not in general solve Eq. (11); for s = 2, the allocation from Eq. (12) does not satisfy Eq. (11) (LHS ≈ 29.9, RHS ≈ 5.1). Thus the closed-form result (12) is unsupported even under the paper's own noncentrality definition.","section":"Appendix 4"}],"minor_comments":[{"comment":"The text repeatedly uses 'close-form' where 'closed-form' is meant; this should be corrected throughout.","section":"§2.3 heading"},{"comment":"Power is defined as the probability that either |Z1| or |Z2| exceeds a cutoff, but Section 2.4 defines estimated power as the minimum of the two empirical rejection proportions. These are different quantities and the inconsistency should be clarified.","section":"§2.3, 'Power definition'"},{"comment":"The caption contains the typo 'Trail parameter estimation'; this should be 'Trial parameter estimation'.","section":"Table 1 caption"},{"comment":"The text says the initial sample size in the binary search is 20, but Section 2.4 Step 1 describes searching from an initial small N0; the connection between these would be clearer if the default N0 were stated.","section":"§3.3, 'Simulation process'"}],"recommendation":"reject","confidential_remarks":"The two algebraic errors in Eq. (7) and Appendix 4 are independent and both invalidate the central design results; the simulation and real-data sections would need to be redone. In addition, the conceptual issue about inter-arm correlations in randomized parallel-arm trials may require reframing the intended design setting. These are substantial, load-bearing problems rather than presentation issues, so I recommend rejection rather than minor or major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is clearly written and the authors have done real work: an R package, simulations, a PDX application, and a K-substudy extension. The generalized Dunnett critical-value computation for a supplied correlation matrix is straightforward and may be useful. But the core derivation is wrong.\n\nSection 2.3 defines the Wald noncentrality as W_AB = delta_AB^2 / Var(mean difference). From the model, Var = (sigma^2/N)[1/p_AB + 1/p_A - 2 rho_AB,A/sqrt(p_AB p_A)], so the correct W is proportional to the reciprocal of that bracket. Equation (7) and Appendix 2 multiply by the bracket instead of dividing. The same inversion appears in the K-substudy version, Eq (16). Everything downstream — the equating step (10), the closed-form allocation (12), the numerical optimization in Appendix 3, and the sample-size search — uses these inverted expressions. The stress-test note is right, and the error is not cosmetic: for s ≠ 1, the Eq (12) allocations do not equalize the correctly defined W_AB and W_B, so they solve a different problem and can reduce power.\n\nThere is a second, independent problem. The paper assumes nonzero correlations rho_AB,A and rho_AB,B between sample means from different arms in a randomized parallel-arm trial. In that setting, patients in different arms are disjoint, so the sample means are independent; the only correlation comes from the shared control. The Appendix tries to justify nonzero correlation via potential outcomes, but randomization makes the groups independent samples. The PDX data are paired, which is a different design. The paper never explains how a trial with genuinely correlated arm means would be randomized or analyzed. If those correlations are zero, Eq (2) collapses to the classical Dunnett correlation and the claimed generalization disappears.\n\nWhat is salvageable: the generalized Dunnett critical value and the error-metric control, if the correlation matrix is treated as an external input from paired preclinical data. But the main design contribution — the power-maximizing allocation and sample-size recommendations — is invalid as stated. This deserves a serious referee to catch the algebra and the design-setting confusion, but I would expect rejection or major revision in current form. I would not cite it yet.","headline":"The paper's central power and allocation derivation inverts the variance denominator, so the claimed optimal designs and sample sizes don't follow; the false-positive-control part may be salvageable, but the main design contribution is invalid as stated.","tokens_in":27037,"tokens_out":2421,"would_cite":false,"duration_ms":28175,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62K05","62J15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A generalized Dunnett's procedure plus a synergy parameter yields optimal arm allocation and false-positive control in combination platform trials.","keywords":["combination therapy","platform trial","multiplicity adjustment","synergy modeling","sample size optimization","Dunnett's procedure","false positive control","translational research"],"falsifier":"Estimate $\\rho_{AB,A}$ and $\\rho_{AB,B}$ from a conventional randomized three-arm trial under the global null: if the observed correlations between arm-level means cluster around zero to within sampling error, Equation (2) reduces to the classical Dunnett correlation and the claimed need for a generalized procedure is not supported by the data.","tokens_in":25862,"feed_emoji":"🧪","tokens_out":8269,"duration_ms":76323,"temperature":0.7,"pith_summary":"This paper proposes a statistical design framework for early-phase platform trials that test a combination therapy (A+B) against both a monotherapy (B) and a shared standard-of-care control (A). The authors extend the classical multiple-comparison adjustment to handle the correlation between test statistics that arises when arms share treatment components, and they show how to control several false-positive metrics (FWER, FMER, MSFP) by solving for a single adjusted critical value. They further derive optimal allocation ratios that maximize the minimum power of the two comparisons, and, under an independence assumption between the combination and control arms, obtain a closed-form allocation that depends only on a synergy parameter $s$. The paper also provides a Monte Carlo and binary-search pipeline for determining the minimal total sample size, and demonstrates the whole workflow on patient-derived xenograft data using an accompanying R package.","feed_headline":"Synergy parameter drives optimal arm allocation in combo trials","feed_subtitle":"A generalized Dunnett procedure controls multiple error rates and converts a synergy estimate into a leaner design.","key_machinery":"The core device is the correlation formula (Eq. 2) for the two test statistics, which expresses the correlation between the combination-versus-control and monotherapy-versus-control comparisons in terms of inter-arm endpoint correlations and sample sizes. This correlation replaces the classical shared-control-only correlation in the joint null distribution, and the same formula is embedded in the Wald noncentrality parameters (Eq. 7) that define power. Equating the two noncentrality parameters under the assumption $\\rho_{AB,A}=0$ yields the closed-form optimal allocation (Eq. 12), a result that depends only on the synergy parameter $s$.","core_discovery":"The central discovery is that the design of a combination platform trial can be driven by a single synergy parameter $s$, defined by $\\delta_{AB}=s\\delta_B$, once the correlation between test statistics is correctly specified. The authors show that the correlation between the two $z$-statistics is governed by inter-arm endpoint correlations and sample sizes (Eq. 2), and that replacing the classical shared-control-only correlation with this full correlation yields a generalized Dunnett's procedure that controls FWER, FMER, and MSFP at user-chosen targets. For power, the Wald noncentrality parameters for the two comparisons are cast in closed form (Eq. 7), and maximizing the minimum of the two leads to a closed-form allocation (Eq. 12), namely $p_A^*=(\\sqrt{s+1}-1)/s$, $p_B^*=(s+1-\\sqrt{s+1})/(s+1)$, and $p_{AB}^*=(s+1-\\sqrt{s+1})/(s(s+1))$, under the assumption $\\rho_{AB,A}=0$. The authors validate via simulation that the procedure controls error rates and that higher synergy reduces the required sample size while shifting allocation from the combination arm to the monotherapy arm. A real-data analysis using patient-derived xenograft models illustrates the full pipeline.","pith_inferences":["The closed-form allocation could be tested in a simulation where $\\rho_{AB,A}$ is small but nonzero; the paper's formula would still be applied, and a power comparison against the numerical optimum would show how much efficiency is lost.","A natural extension is to treat $s$, $\\delta$, and the inter-arm correlations as uncertain priors rather than point estimates; the design pipeline could then report robust allocations that hedge against misspecified preclinical translation.","The max-min power criterion treats both hypotheses equally, but regulatory priorities often emphasize the combination hypothesis; re-weighting the objective would shift the closed-form solution, and sensitivity analysis could reveal whether the allocation is stable.","The same correlation machinery could apply to trials with more than two active components, but the closed-form allocation would likely disappear because the number of equality constraints exceeds the degrees of freedom."],"forward_implications":["Combination-trial designers can compute power-maximizing allocation ratios directly from a synergy estimate without numerical optimization when the combination and control arms are uncorrelated.","The generalized Dunnett procedure controls FWER, FMER, and MSFP at user-chosen levels across the simulated range of arm correlations, avoiding the over-conservatism of Bonferroni and Holm and the mis-specified correlation of the classical Dunnett test.","The pipeline returns a minimal total sample size for a target power while maintaining the chosen false-positive control, enabling pre-trial resource planning.","Preclinical patient-derived xenograft data can be plugged into the design pipeline to estimate effect sizes, synergy, and inter-arm correlations, making early-phase trials more informative.","The framework extends to $K$ substudies within one platform, each testing one monotherapy and its combination against a common control."],"supporting_citations":[{"why":"Supplies the classical shared-control multiple comparison procedure that the paper generalizes to overlapping arms.","marker":"[7]"},{"why":"Provides the patient-derived xenograft dataset with paired endpoints used to estimate effect sizes, synergy, and arm correlations.","marker":"[24]"},{"why":"Defines the false-positive metrics and open-entry platform setting that motivate FMER and MSFP control.","marker":"[13]"},{"why":"Gives recommendations on multiple testing adjustment in multi-arm trials with a shared control, the context the paper extends.","marker":"[11]"},{"why":"Supports the max-min power criterion for optimizing allocation across arms.","marker":"[22]"},{"why":"Provides the potential-outcomes justification used in the appendix for nonzero correlations between sample means.","marker":"[39]"},{"why":"Supplies context on allocation and treatment selection in multi-arm multi-stage designs.","marker":"[15]"},{"why":"Describes multi-arm group-sequential designs that the paper positions against and cites for future adaptive extensions.","marker":"[9]"}],"fun_headline_variants":["Synergy parameter shapes combo-trial design","One synergy number sets combo trial arm sizes","Generalized Dunnett tames combo trial errors","Synergy-based allocation for combo platform trials","Single parameter guides optimal combo trial design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the endpoint correlations between the combination arm and the control ($\\rho_{AB,A}$) and between the combination and monotherapy ($\\rho_{AB,B}$) are real, nonzero quantities that can be estimated from preclinical paired data and will persist in a randomized clinical trial; if these correlations are zero by design, the generalized correlation formula and the closed-form allocation collapse.","fun_headline_variants_meta":{"raw":{"variants":["Synergy parameter shapes combo-trial design","One synergy number sets combo trial arm sizes","Generalized Dunnett tames combo trial errors","Synergy-based allocation for combo platform trials","Single parameter guides optimal combo trial design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000432,"raw_usage":{"total_tokens":2236,"prompt_tokens":1009,"completion_tokens":1227,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":1161}},"tokens_in":625,"tokens_out":1227,"duration_ms":9259,"temperature":1.0,"reasoning_tokens":1161,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:11:20.681698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate $\\rho_{AB,A}$ and $\\rho_{AB,B}$ from a conventional randomized three-arm trial under the global null: if the observed correlations between arm-level means cluster around zero to within sampling error, Equation (2) reduces to the classical Dunnett correlation and the claimed need for a generalized procedure is not supported by the data.","supporting_citations":[],"review_version":1}