{"id":"ef41f88d-a106-451b-89a4-9eb04edccc69","arxiv_id":"2411.13692","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RaBIt extends Chen et al.'s D2 basket trial to unequal basket sizes and effects, maintaining type 1 error with a weighted pooled test and cutting expected trial duration when allocation matches accrual rates.","lead":"This paper generalizes a two-stage randomized basket trial so that different disease groups can have different sample sizes and expected effects, and shows that matching enrollment to real-world accrual speeds can shorten trial duration with almost no loss of statistical power. It derives formulas for power and type 1 error and applies the design to a possible psilocybin trial for OCD, body dysmorphic disorder, and anorexia nervosa.","discovery_kind":"extension","skeptic_critique":null,"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RaBIt, an extension of Chen et al.'s D2 randomized basket trial design. The key generalization is to allow baskets to have different planned sample sizes and different anticipated effect sizes. The design prunes baskets at an interim analysis based on a common threshold and then combines the remaining baskets using a weighted Stouffer statistic, with weights determined by the initial allocation proportions conditionally on the set of baskets retained. The authors derive an expression for the overall type 1 error as a sum over all possible retention sets, solve numerically for the final critical value alpha*, and derive the corresponding power formula. They validate the implementation by reproducing Chen et al.'s power values for equal-sized baskets, examine how alpha* and power vary with allocation imbalance (measured by a Gini impurity), and compute expected trial duration and sample size under constant accrual rates. In a worked example, allocating baskets proportionally to accrual is reported to shorten expected duration by about 17.5 months with a power loss of 0.25 percentage points.","tokens_in":14877,"tokens_out":23291,"duration_ms":216974,"significance":"The statistical derivation appears sound and offers a useful, practical extension of a published confirmatory basket trial design. The paper ships code and validates against Chen et al.'s published results, which is a concrete strength that supports reproducibility. The design is relevant to mental health and other settings where basket accrual rates and anticipated effect sizes are heterogeneous. If the duration result is robust, it has clear logistical value. The main caveat concerns the interim-timing assumption underlying the duration comparisons, which needs to be surfaced and tested.","major_comments":[{"comment":"The headline duration saving of about 17.5 months is computed under the 'fastest possible, though impractical' assumption stated in Appendix B.1, namely that each basket's interim analysis is performed as soon as that basket reaches its interim target sample size. This assumption is not disclosed in the abstract or in Section 3.5, where the duration reduction is presented as a design benefit. Since a conventional trial would conduct a single interim analysis only after all baskets complete stage 1 accrual, and the authors themselves note this simpler strategy 'will increase the trial duration', the reported saving may not be realized in practice. Please provide expected durations under the single-interim-time model as well, or prominently qualify the claim in the abstract and Section 3.5.","section":"Section 3.5 / Table 3 / Abstract"},{"comment":"The displayed event for pruned baskets is written as \\cap_{l \\notin id(m)} Y_{l1} > Z_{1-\\alpha_t}, which would require the pruned baskets to also exceed the interim threshold. This is inconsistent with the factorization in equation (5), which multiplies by (1-\\alpha_t)^{K-|id(m)|} for those baskets, and with the corresponding event in equation (11), where pruned baskets satisfy Y_{j1} < Z_{1-\\alpha_t}. The inequality in equation (4) should be corrected; as printed, the event is empty whenever any basket is pruned, which would make the subsequent formula unintelligible.","section":"Equation (4)"},{"comment":"The sentence 'the product of baskets accurately getting pruned away (let there be R of them) and baskets inaccurately getting pruned away' is reversed relative to the formula that follows. The product over id(g)\\id(j) corresponds to active baskets that are incorrectly pruned, while (1-\\alpha_t)^R is the contribution of inactive baskets that are correctly pruned. Please reword the explanation so that the text matches the displayed expression.","section":"Section 2.3, power decomposition"}],"minor_comments":[{"comment":"The rendering of corr(Y_{i2}, V_m) as 'w_i qP_m i=1 w^2_i' is garbled; it should read w_i / sqrt(\\sum_{i\\in id(m)} w_i^2). Please fix the typesetting.","section":"Equation (8)"},{"comment":"There is a missing word: 'consistent the prior methods' should be 'consistent with the prior methods'.","section":"Abstract"},{"comment":"The statement that power values 'only deviate ±0.3%' is not supported by Table 1, where all absolute differences are on the order of 10^-4 (i.e., roughly 0.01 percentage points). Please report the actual maximum deviation.","section":"Section 3.1"},{"comment":"The text defines t as the information time but writes N \\cdot p_i \\cdot t_i for each basket; the subsequent formulas (e.g., equation (7)) use a common t. Clarify that a single information time is assumed for all baskets.","section":"Section 2.1"},{"comment":"The sum over m in equation (10) should state explicitly that terms with m = 0 (no baskets retained) contribute zero probability to the overall type 1 error, since no final test is performed in that case.","section":"Section 2.2 / Equation (10)"},{"comment":"The 'Gini Impurity' used here, 1 - \\sum p_i^2, is not the usual Gini coefficient; higher values indicate more equal allocation. A one-line explanation of the measure's interpretation would help avoid confusion.","section":"Section 2.4 / Figure 2"},{"comment":"The heuristic explanation for why unequal allocation leads to a less stringent alpha* refers to 'interim power' under the alternative, whereas the alpha* calibration is derived under H0 where sample size does not affect the marginal distribution of each interim z-statistic. The heuristic may be confusing and should be reformulated in terms of the correlations in equation (9).","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The statistical methodology is fundamentally sound and well aligned with the journal's scope. The main concern is that the paper's most practically appealing claim, the large reduction in trial duration, is presented in the abstract and Section 3.5 without the caveat that it relies on the 'fastest possible, though impractical' interim-timing model described only in the supplement. A revision should either provide duration results under a conventional single-interim-analysis schedule or prominently qualify the claim. The equation (4) typo, while clearly not intended, must be corrected because it appears in the central derivation. A small simulation study of the type 1 error would further strengthen the paper, but the analytical derivation is consistent and I would not require it as a condition for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: RaBIt does what it claims. It generalizes the D2 randomized basket trial to baskets of different sizes and effect sizes, and it derives analytical type 1 error and power formulas that reproduce Chen et al.'s numbers to within 0.3% in the equal-size case. The weighted Stouffer statistic and the proportional reallocation rule w_i = p_i/(m·p) are the right natural extensions. The correlation calculation between interim and final statistics is correct, and summing over all m is a valid unconditional calculation. The trial-duration analysis, while based on an idealization, supports the practical claim: matching allocation to accrual rates can shorten expected duration substantially at negligible power cost.\n\nSoft spots, in order of real importance. (1) The abstract says the final threshold becomes more stringent when sample allocation is unequal, but the results and discussion say the opposite - it becomes less stringent (alpha* rises from 0.0143 to 0.0192 in the K=2 case). One of those is wrong and the abstract likely is. (2) The duration numbers depend on running each basket's interim as soon as it reaches its target, which the paper itself calls 'fastest possible, though impractical.' The advantage under a real calendar-driven interim is not shown. (3) Equation (4) has a sign error; it is corrected in (5) and (11), but it will confuse readers. (4) The effect-size parameterization (mean = Delta sqrt(n)/4) is nonstandard and deserves a one-line explanation. (5) The code and Shiny app are referenced but were not available for inspection, so I could not verify the unequal-case computations independently.\n\nNone of these are load-bearing. The statistical argument is self-contained and the equal-size validation gives me confidence the unequal-case results are not curve-fitting. The limitations section is honest about what was not varied.\n\nBottom line: send it to review. A competent referee can sort out the abstract, the typo, and the duration caveat without any new derivations. I'd cite it as the natural reference for unequal-size randomized basket trials, and I'd bring it to a reading group focused on master protocols.","headline":"A sound, useful generalization of Chen et al.'s randomized basket trial to unequal baskets; the statistics hold up, but the write-up has a few inconsistencies that need cleaning.","tokens_in":15434,"tokens_out":1939,"would_cite":true,"duration_ms":55143,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62L05","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"RaBIt generalizes randomized basket trials to unequal basket sizes and effect sizes while preserving overall type 1 error, and in a worked example shortens expected trial duration by about 17.5 months at a power loss of roughly 0.0025.","keywords":["randomized basket trial","interim analysis","pooled analysis","type 1 error control","weighted Stouffer's Z","unequal sample sizes","accrual-rate matching","mental health trials"],"falsifier":"Simulate the worked three-basket trial under a misspecified correlation or with non-normal endpoints and check whether the empirical type 1 error matches the nominal $\\alpha = 0.025$; separately, recompute expected duration under the practical policy of a single interim analysis after all baskets finish stage 1. If the empirical type 1 error deviates materially, or the 17-month duration gap closes, the claims are conditional on those assumptions.","tokens_in":14735,"feed_emoji":"📊","tokens_out":9305,"duration_ms":89322,"temperature":0.7,"pith_summary":"RaBIt is a two-stage randomized basket-trial design that removes the equal-basket restriction of earlier pruning-and-pooling designs. Each basket may have its own planned sample size and its own anticipated treatment effect; at an interim analysis, unpromising baskets are pruned, and the surviving baskets are analyzed together with a weighted pooled test. The paper derives analytic expressions for the power and overall type 1 error of this design and calibrates the final test threshold $\\alpha^*$ by solving a sum over all possible pruning configurations. It then shows that when all baskets are equal, RaBIt reproduces the earlier D2 design's power to within $\\pm 0.3$ percentage points. In the motivating three-basket mental-health example, matching basket sizes to accrual rates shrinks the expected trial duration from about 61 to about 44 months, at a power cost of about 0.25 percentage points.","feed_headline":"Matching basket size to accrual cuts trial time by 17 months","feed_subtitle":"A generalized two-stage basket design keeps type 1 error fixed while baskets differ in size and effect.","key_machinery":"The load-bearing object is the weighted pooled statistic $V_m$ and the weights $w_i = p_i/(m \\cdot p)$, which define how a pruned basket's sample is redistributed to the baskets that remain. Because the weights are proportional to the original basket proportions, the final pooled analysis stays aligned with the planned unequal design. The paper's calibration step uses the independent-increments correlation $\\mathrm{corr}(Y_{i1}, Y_{i2}) = \\sqrt{t\\,(m \\cdot p)}$ together with the weight correlation $\\mathrm{corr}(Y_{i2}, V_m) = w_i / \\sqrt{\\sum w_i^2}$ to express each configuration's rejection probability, then solves numerically for the final threshold $\\alpha^*$ from $\\alpha = \\sum_{m\\in M} \\Pr_{H_0}(V_m \\mid \\alpha^*, \\alpha_t, m)$. This converts the equal-basket combinatorics of the earlier design into a simple sum over pruning configurations.","core_discovery":"The central claim is that a randomized basket trial can prune and pool baskets of different sizes and different effect sizes without inflating the overall type 1 error. For each possible set of baskets that survives the interim, the final test statistic is the weighted Stouffer combination $V_m = (\\sum_{i \\in \\mathrm{id}(m)} w_i Y_{i2}) / \\sqrt{\\sum_{i \\in \\mathrm{id}(m)} w_i^2}$, with weights $w_i = p_i/(m \\cdot p)$ that reallocate the sample mass of pruned baskets in proportion to each surviving basket's original size. The paper derives the correlation between the interim statistic and this final statistic, and obtains $\\alpha^*$ by requiring the sum of rejection probabilities over all pruning configurations to equal the nominal $\\alpha$. Under this calibration, power also has a closed-form sum. The authors report that equal baskets recover the D2 design's powers almost exactly, while unequal allocation makes the final threshold less stringent at a small power cost; in the worked example, proportional-to-accrual allocation reduces expected duration by roughly 17.5 months at a power loss near 0.0025.","pith_inferences":["The 17-month duration saving is tied to the paper's 'fastest possible, though impractical' assumption that the interim analysis is run as soon as each basket reaches its target; under the more realistic policy of one interim analysis after all baskets finish stage 1, the duration advantage could shrink or disappear.","The $\\alpha^*$ calibration assumes normally distributed interim and final statistics with known variance and the independent-increments correlation of equation (7); a misspecified correlation, or non-normal endpoints, would require a simulation check before the thresholds could be trusted in practice.","The same weighting idea could be carried to umbrella or platform designs where pruning decisions are made per subgroup and final inference is shared through a common control; the paper does not develop that direction.","A natural stress test would compare RaBIt's frequentist operating characteristics with Bayesian hierarchical basket designs under prior misspecification, since RaBIt deliberately avoids information sharing between baskets."],"forward_implications":["A phase 3 basket trial can plan basket sizes to match expected accrual without losing type 1 error control; in the paper's three-basket example, this reduces expected duration from about 61 to about 44 months with a power difference of about 0.0025.","More unequal basket allocation makes the final threshold $\\alpha^*$ less stringent, for example $\\alpha^* = 0.0100$ at equal sizes versus $0.0152$ at the most unequal allocation tested for three baskets, while lowering power modestly from 0.879 to 0.837.","If effect sizes are unequal but their average is fixed, concentrating the larger effect in one basket increases overall power; for average effect 0.5, increasing one basket's effect from 0.5 to 1.1 raises power from 0.879 to 0.969.","When baskets are equal, the generalized formulas reproduce the earlier D2 design's power within $\\pm 0.3\\%$, so RaBIt is a backward-compatible extension.","A frequentist, prior-free randomized basket design is available for confirmatory mental-health trials where accrual differs by indication."],"supporting_citations":[{"why":"Supplies the D2 pruning-and-pooling design that RaBIt generalizes and the baseline code whose outputs are validated.","marker":"Chen et al. (2016)"},{"why":"Provides the weighted Stouffer's Z formula used to construct the final pooled test statistic $V_m$.","marker":"Stouffer et al. (1949)"},{"why":"Defines master protocols as the motivating framework for testing one treatment across multiple diseases.","marker":"Woodcock and LaVange (2017)"},{"why":"Reviews basket, umbrella, and platform trials, positioning the efficiency goal of the proposed design.","marker":"Renfro and Sargent (2017)"},{"why":"Documents late-phase randomized basket trial practice that the paper targets for generalization.","marker":"Kasim et al. (2023)"}],"fun_headline_variants":["Generalized basket design cuts trial time by 17.5 months","Type 1 error fixed while basket sizes differ","Prune and pool: basket trial design with flexible accrual","RaBIt generalizes basket trials for speed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the interim and final test statistics follow the normal, known-variance, independent-increments correlation structure in equation (7); the trial-duration numbers also rely on constant accrual and on running interim analyses as soon as each basket reaches its target, a strategy the paper calls 'fastest possible, though impractical.'","fun_headline_variants_meta":{"raw":{"variants":["Generalized basket design cuts trial time by 17.5 months","Type 1 error fixed while basket sizes differ","Prune and pool: basket trial design with flexible accrual","RaBIt generalizes basket trials for speed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000578,"raw_usage":{"total_tokens":2738,"prompt_tokens":969,"completion_tokens":1769,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":1703}},"tokens_in":585,"tokens_out":1769,"duration_ms":14407,"temperature":1.0,"reasoning_tokens":1703,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:00:12.840717+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the worked three-basket trial under a misspecified correlation or with non-normal endpoints and check whether the empirical type 1 error matches the nominal $\\alpha = 0.025$; separately, recompute expected duration under the practical policy of a single interim analysis after all baskets finish stage 1. If the empirical type 1 error deviates materially, or the 17-month duration gap closes, the claims are conditional on those assumptions.","supporting_citations":[],"review_version":1}