{"id":"9971b18c-75c0-44a6-b384-6294073fb5ed","arxiv_id":"2411.15567","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper extends Japan's MHLW regional-consistency criteria to two pivotal multi-regional clinical trials and derives regional sample-size formulas with simulation validation.","lead":"This paper develops formulas for how many patients a region needs to enroll when it joins two large global drug trials, so the region can later show its results match the overall drug effect. Regulators in Japan and China often require such regional consistency evidence, and no formal tool previously existed for the two-trial case.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 4's pooling identity (Remark 9) is exact only for equal randomization ratios; Tables 5–6 apply it with r(1)=1, r(2)=2, so the central CP formula is used in a regime where its derivation is algebraically misspecified.","rationale":"The reader's declared weakest assumption is the fixed-effects common-effect model. That is a legitimate scope limitation, and the Discussion explicitly cautions against pooling when covariate distributional shifts differ across studies, so it is a modeling assumption rather than an internal flaw. The more damaging issue is internal: the derivation of the paper's central formula (13) uses a linear pooling identity that is false in the unequal-randomization-ratio setting the paper itself simulates. This changes the conditional mean and variance used in the integrand, not merely a philosophical detail. The B=100,000 simulations are reassuring but cannot replace a correct derivation; near-nominal CP at selected f-values could reflect cancellation or small deviations at those specific design points. The proposed check directly compares the true pooled-estimator CP with Eq. (13) in an unequal-ratio table row. If the check shows only sub-percentage-point differences over the tested grid, the paper is acceptable after adding a condition or qualification to Remark 9; if the difference is material, the sample-size solutions for unequal ratios need revision. The reader already returned a CONDITIONAL verdict, and this concern reinforces rather than changes that verdict.","tokens_in":21056,"tokens_out":6621,"duration_ms":59555,"concrete_test":"Use the Table 6 row (1−β1=0.8, 1−β2=0.9, d=1, σ=4, r(1)=1, r(2)=2) with reported regional fractions f(1)=0.114, f(2)=0.121. Compute the true consistency probability of the estimator defined in Eq. (11) by deriving the exact conditional distribution of D_k,pool − πD_pool given D(1),D(2), using treatment-arm weights w_t(s)=N_t(s,k)/(N_t(1,k)+N_t(2,k)) and control-arm weights w_c(s)=N_c(s,k)/(N_c(1,k)+N_c(2,k)), where N_t(s,k)=f(s)·r(s)/(1+r(s))·N(s) and N_c(s,k)=f(s)/(1+r(s))·N(s). Evaluate this exact probability by numerical integration or by B=100,000 simulation under the same parameter configuration, and compare it with the nominal 0.80 and with the value returned by Eq. (13). If the exact CP differs by more than about 0.5 percentage points, Eq. (13) is misspecified in the unequal-ratio regime and Tables 5–6 cannot validate the method as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 4 (Eq. 13) and Remark 9 rely on the decomposition D_k,pool = w(1)D_k(1) + w(2)D_k(2) and D_pool = w(1)D(1) + w(2)D(2), with w(s)=N(s)/(N(1)+N(2)). Under the actual pooled definitions in Eq. (11), this equality is exact only when the two studies share the same randomization ratio. The pooled treatment mean is a weighted average of the study-specific treatment means with weights N_t(s,k)/Σ_u N_t(u,k), and the pooled control mean uses weights N_c(s,k)/Σ_u N_c(u,k). These two sets of weights coincide only if N_t(s,k)/Σ_u N_t(u,k) = N_c(s,k)/Σ_u N_c(u,k) for each s, which, given the assumed constant randomization ratio within each study, is equivalent to r(1)=r(2). When r(1)≠r(2), D_k,pool and D_pool are not the simple w-weighted sums used in the proof of Proposition 4, so the conditional mean (1−π)(w(1)D(1)+w(2)D(2)) and the conditional variance feeding into (13) are not the correct ones for the estimator defined in (11). This matters because Tables 5–6 are exactly the unequal-ratio case r(1)=1, r(2)=2. The simulations show CP near nominal, but that only demonstrates that the misspecification was empirically small at the tested f values, not that Eq. (13) evaluates the CP of the pooled estimator. The paper should either restrict the method to r(1)=r(2), or re-derive the formula using arm-specific weights and state the conditions under which the w-weighted approximation is valid.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses regional consistency evaluation and sample size calculation for two multi-regional clinical trials (MRCTs). The authors first review existing criteria for a single MRCT under a fixed-effects model and present unified approximations for the MHLW type I and type II consistency probabilities (Propositions 1–3). They then extend these criteria to the setting where a region participates in two pivotal MRCTs, proposing pooled-data criteria (12) and (16). Proposition 4 gives an approximate conditional consistency probability for criterion (12) as a double integral; Propositions 5 and 6 give the corresponding results for the all-regions criterion and for binary endpoints. Regional sample size fractions are solved numerically, with a closed-form optimality condition (14) for minimizing the combined regional sample size. Simulations with B = 100,000 replications and a hypothetical example illustrate the methods, and an R package is provided.","tokens_in":21288,"tokens_out":18285,"duration_ms":138631,"significance":"The two-MRCT extension fills a genuine practical gap, as regions sometimes need to enroll across both pivotal trials. The paper provides a clear framework, ready-to-use numerical integration schemes, extensive simulations, and an R package; the one-MRCT results are standard and the empirical CP values are close to nominal in the reported grids. The key novelty, however, rests on Proposition 4, and that derivation contains an algebraic misspecification: the identities in Remark 9 hold only under equal randomization ratios and equal regional fractions across studies. For the general case (e.g., Tables 5–6 with r(1)=1, r(2)=2), Eq. (13) does not evaluate the consistency probability of the pooled estimator defined in Eq. (11). This must be corrected before the method can be recommended for general use.","major_comments":[{"comment":"The identities D_{k,pool}=w^{(1)}D_k^{(1)}+w^{(2)}D_k^{(2)} and D_{pool}=w^{(1)}D^{(1)}+w^{(2)}D^{(2)} stated in Remark 9 and used in the proof of Proposition 4 are not identities for the estimators defined in Eq. (11). From (11), D_{pool} is a weighted average with treatment-arm weights N_t^{(s)}/(N_t^{(1)}+N_t^{(2)}) and control-arm weights N_c^{(s)}/(N_c^{(1)}+N_c^{(2)}); these equal w^{(s)}=N^{(s)}/(N^{(1)}+N^{(2)}) only if r^{(1)}=r^{(2)}. Similarly, D_{k,pool} has weights proportional to f_k^{(s)}N_t^{(s)} and f_k^{(s)}N_c^{(s)}, which equal w^{(s)} only if, additionally, f_k^{(1)}=f_k^{(2)}. Consequently, the conditional mean (1−π)(w^{(1)}D^{(1)}+w^{(2)}D^{(2)}) and the variance terms ((f_k^{(1)})^{-1}−1)(w^{(1)}σ_d^{(1)})^2 + ((f_k^{(2)})^{-1}−1)(w^{(2)}σ_d^{(2)})^2 in Eq. (13) are not the moments of D_{k,pool}−πD_{pool} under the definition (11). Tables 5–6 and S3–S4 use exactly the unequal-ratio case r^{(1)}=1, r^{(2)}=2 (and unequal f_k^{(s)}), so the close agreement of empirical CP does not validate the derivation; it only indicates that the misspecification is numerically small in those scenarios. The authors should either restrict Propositions 4–6 to settings where the weighting identities hold (r^{(1)}=r^{(2)} and f_k^{(1)}=f_k^{(2)}) or re-derive the consistency probability and sample-size formulas using the arm-specific weights in (11).","section":"2.3, Remark 9 and Appendix A.3 (Eq. 13)"},{"comment":"The minimization condition (14) is a consequence of the variance expression in (13). Because that variance is not the variance of the actual pooled estimator when the weights differ, the reported 'combined sample size minimized' solutions in Tables 5–6 and S3–S4 are not necessarily optimal for the estimator in (11). The authors should re-derive the optimality condition from the correct conditional variance, or state explicitly that (14) optimizes the approximate variance in the w-weighted model rather than the actual pooled estimator.","section":"2.3, Eq. (14) and Remark 6"},{"comment":"The event in (18) compares w^{(1)}p̂_{t,k}^{(1)}+w^{(2)}p̂_{t,k}^{(2)} with w^{(1)}p̂_{c,k}^{(1)}+w^{(2)}p̂_{c,k}^{(2)}, which is the w-weighted combination of study-specific regional differences. This is not the same as D_{k,pool} ≥ 0 defined by the pooled estimator in (11) unless the weights coincide as in the first comment. The same correction is needed for the binary-response exact calculation; otherwise Proposition 6 evaluates a different event from the LHS of (16).","section":"2.3, Proposition 6 (Eq. 18)"}],"minor_comments":[{"comment":"The acronym 'MHLW' is inconsistently spaced as 'MHL W' in the abstract, introduction, and elsewhere; please make it consistent.","section":"Throughout"},{"comment":"There is a typo: 'respetively' should be 'respectively'.","section":"Example 2"},{"comment":"The line 'p(c,1) = p(c,1) = 0.8' should read 'p(c,1) = p(c,2) = 0.8'.","section":"Example 4"},{"comment":"In the sentence defining the binomial distributions, 'b(s)_k ∼ Bin(f(s)_k N(c,s), p(t,s))' should use p(c,s) rather than p(t,s).","section":"Proposition 6"},{"comment":"The claim that the approximation error in (6) is O(n^{-1/2}) is not demonstrated in the Appendix; the proof only shows the asymptotic approximation, not the stated rate. Either provide a proof of the rate or soften the claim to 'the approximation is asymptotically justified'.","section":"Remark 2"},{"comment":"The sentence 'The global enrollment of Study 1 and Study 2 is sequential' should be 'are sequential'.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The pooling-weight misspecification is the main technical concern. It is fixable either by restricting the scope (which would reduce the contribution, since the unequal-ratio case is explicitly featured) or by re-deriving the formulas with the correct arm-specific weights. The numerical agreement in the tested tables suggests the discrepancy may often be small, but the manuscript's central claim is not established for the general case as written. The paper is otherwise clearly organized, and the R package and extensive simulations are valuable assets."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know: this paper genuinely extends the MHLW regional-consistency criteria from a single MRCT to the two-pivotal-MRCT setting, and the simulation work is careful enough that the empirical claims are credible. The closed-form CP integrals (13), (15), (17), the minimal-combined-sample-size condition (14), and the binary-endpoint expression (18) are new relative to the cited single-MRCT literature. No fitted parameters, no circularity; code and an R package are provided.\n\nThe soft spot is the pooling identity used in Remark 9 and the proof of Proposition 4. The paper writes D_{k,pool} = w^(1)D_k^(1) + w^(2)D_k^(2) and D_pool = w^(1)D^(1) + w^(2)D^(2) with w^(s)=N^(s)/(N^(1)+N^(2)). Under the definitions in (11), that equality is not exact. It requires equal randomization ratios across the two studies and, for the regional estimator, equal regional fractions f_k^(1)=f_k^(2). Tables 5 and 6 use r^(1)=1, r^(2)=2, with unequal f_k's, so the central formula (13) is used in a regime where its derivation is an approximation. The simulations show the achieved CP stays within about a percentage point of nominal, so the misspecification looks benign in the tested grid, but the paper should state the exact conditions or justify the approximation. This is a curable revision, not a fatal flaw.\n\nThe other main caveat is the fixed-effects common-effect assumption, which is load-bearing and which the authors acknowledge in the Discussion, where they note pooled assessment is inappropriate when covariate-shift patterns differ across studies. That is honest and correctly scoped. Minor typos: in Proposition 6 the second binomial has p(t,s) where it should be p(c,s), and Example 4 repeats p(c,1). Also the notation around S(z_{1-α}) in Proposition 3 is a little compressed. These are cosmetic.\n\nFor a biostatistician planning regional enrollment in a two-MRCT program, this is a useful design tool and I would cite it. It deserves a serious referee; the main request should be to make the approximate nature of the pooling identity explicit and, if possible, derive the correct weights or bound the error.","headline":"Genuine extension of MHLW regional-consistency methods to two-pivotal-MRCT planning, with solid simulations, but the central pooling identity is an unstated approximation that needs fixing.","tokens_in":21995,"tokens_out":7710,"would_cite":true,"duration_ms":57817,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62K05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper derives double-integral formulas for regional consistency after two pivotal MRCTs are both significant, and shows these formulas can be solved to set regional sample fractions.","keywords":["multi-regional clinical trials","regional consistency evaluation","sample size calculation","fixed effects model","MHLW criteria","consistency probability","pooled analysis","binary endpoints"],"falsifier":"Run the two-trial simulation with a binary endpoint at the small end of the paper's grid, such as $p^{(c)}=0.8$, $d=0.2$, and $N^{(s)}=118$; solve $f_k$ from the normal formula (13), and compare the empirical conditional consistency probability over many replications to the nominal 80%. An error larger than simulation noise would show the normal approximation fails there. Also simulate with truly different regional effects, for instance region 1 having effect 0.5 while others have 1.0, and check whether the solved fractions still deliver the claimed probability, which would expose the fixed-effects assumption.","tokens_in":20678,"feed_emoji":"📊","tokens_out":9862,"duration_ms":84912,"temperature":0.7,"pith_summary":"Drug registration often requires two pivotal multi-regional trials, and a region may need to join both to reach a required subpopulation. This paper asks how a region should split its enrollment across the two trials so that, once both trials show an overall effect, the region's pooled estimate credibly matches the pooled overall estimate. Under the fixed-effects model with a single common treatment effect, it derives explicit approximations for the probability of regional consistency, expressed as double integrals in the two standardized trial statistics. Solving those integrals gives the regional sample fractions in each trial that attain a desired consistency probability, and the paper's simulations show the empirical probability lands within about one percentage point of the target. The practical payoff is a design tool, implemented in an R package, for regions that must join both pivotal MRCTs.","feed_headline":"Pooling two pivotal trials shrinks the regional sample needed","feed_subtitle":"New formulas split a region's enrollment across two trials so the chance of showing consistency hits the target.","key_machinery":"The machinery is the joint asymptotic normality of the pooled regional and overall estimators, combined with conditioning on both study test statistics exceeding $z_{1-\\alpha}$. The paper forms $D_{k,\\mathrm{pool}}=w^{(1)}D_k^{(1)}+w^{(2)}D_k^{(2)}$ and $D_{\\mathrm{pool}}=w^{(1)}D^{(1)}+w^{(2)}D^{(2)}$, then uses the identity $\\sigma_d^{2(s)} = d^2_{(s)}/(z_{1-\\alpha}+z_{1-\\beta_s})^2$ to turn each rejection event $T^{(s)} > z_{1-\\alpha}$ into $Z^{(s)} > -z_{1-\\beta_s}$. That converts the conditional consistency probability into a double integral with no unknown parameters beyond design inputs, and the required regional fractions depend on $\\alpha$, the powers, effect sizes, variances, and randomization ratios, but not on the total sample sizes $N^{(s)}$.","core_discovery":"The central result is Proposition 4. For two independent pivotal trials with treatment differences $d^{(s)}>0$, the authors approximate the conditional consistency probability $$\\Pr\\!\\left(D_{k,\\mathrm{pool}} \\ge \\pi D_{\\mathrm{pool}} \\mid $T^{{(1)}}$ > z_{1-\\$\\alpha$},\\, $T^{{(2)}}$ > z_{1-\\$\\alpha$}\\right)$$ by a double integral over the standardized trial statistics $u,v$, whose integrand is a normal CDF and whose normalizing denominator is $(1-\\beta_1)(1-\\beta_2)$, as in equation (13). Because this probability is monotone in the combined fraction quantity $\\zeta$, a planner can solve for regional fractions $f_k^{(1)}$ and $f_k^{(2)}$ to hit a target consistency probability, and the combined regional sample size is minimized when the fraction ratio obeys equation (14). The same conditioning device is applied to the second MHLW criterion, where all regions must show the same trend, yielding Proposition 5, and exact binomial versions are given for binary endpoints in Propositions 3 and 6. The paper reports that with these solved fractions the empirical consistency probability tracks the nominal 80% to within roughly one percent across continuous and binary response settings.","pith_inferences":["The conditioning device generalizes naturally: if more than two pivotal trials must all be significant, the same construction gives a higher-dimensional integral, and the equal-fraction optimum under homogeneity should continue to hold, though this is not stated in the paper.","The binary-endpoint correction suggests that the normal approximation should be stress-tested for time-to-event endpoints, where test statistics are not normal in moderate samples; the paper flags this as future work.","A sponsor facing enrollment constraints could replace the minimum-total-sample objective with other constraints, such as a cap on one trial's fraction, because every consistency-probability level has a curve of feasible fraction pairs even though the paper only solves the minimum-size objective.","The paper's own caution implies a boundary: if region-specific covariate distributions shift differently in the two studies, the pooled estimate is not a valid consistency target, and an extension could quantify the bias by modeling the shift explicitly."],"forward_implications":["If formula (13) is right, a region enrolling in both trials no longer needs to carry a large share of either trial; in the paper's worked example, an 80% consistency probability is met with only 10.9% of each study's sample under equal fractions.","Because many fraction pairs $(f_k^{(1)}, f_k^{(2)})$ achieve the same consistency probability, a planner can trade enrollment between studies, such as (8%, 17.4%) or (9%, 14.1%) in the example, while the minimum total regional size pins down the ratio in (14).","The pooled criterion is substantially more efficient than checking each trial separately: in Remark 9, the required fraction drops from 46.6% per trial to 15.4% once the two trials are pooled and both are significant.","For binary endpoints, the normal double-integral approximation under-sizes the region, and the exact binomial formulas are needed; Example 4 shows $f_1 = 6.0\\%$ rather than the normal-approximation value of 4.4%.","The methods come with an R package, so the fraction solving and consistency-probability evaluation are directly usable in practice."],"supporting_citations":[{"why":"Supplies the two consistency criteria and the regional-fraction framework that this paper extends to two pooled trials.","marker":"MHLW (2007)"},{"why":"Gives the single-trial normal approximation and Japanese-patient sample-size method that Proposition 4 generalizes.","marker":"Ikeda and Bretz (2010)"},{"why":"Provides the method for determining a region's sample size in one MRCT, the baseline for the one-trial consistency-probability formula.","marker":"Ko et al., 2010"},{"why":"Provides the fixed-effects decision rules and regional sample-size planning that underpin Section 2.1.","marker":"Chen et al. (2012)"},{"why":"Supplies the exact binomial evaluation of the all-regions-positive criterion that Propositions 3 and 6 adopt.","marker":"Homma (2023)"},{"why":"Sets the regulatory requirement of two adequate and well-controlled trials that motivates the two-MRCT design.","marker":"FDA (2019)"}],"fun_headline_variants":["Two-trial regional consistency: sample size formulas hit 80% target","Pooling MRCTs cuts regional enrollment while matching consistency odds","New formulas for regional sample size under dual pivotal trials","R package solves regional sample size for two MRCTs with exact control","Consistency probability pinned: two MRCTs need less regional data per trial"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume a single true treatment effect shared by all regions and both trials, and that the test statistics are close enough to normal for the double integral to hold.","fun_headline_variants_meta":{"raw":{"variants":["Two-trial regional consistency: sample size formulas hit 80% target","Pooling MRCTs cuts regional enrollment while matching consistency odds","New formulas for regional sample size under dual pivotal trials","R package solves regional sample size for two MRCTs with exact control","Consistency probability pinned: two MRCTs need less regional data per trial"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00036,"raw_usage":{"total_tokens":1963,"prompt_tokens":977,"completion_tokens":986,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":895}},"tokens_in":593,"tokens_out":986,"duration_ms":7535,"temperature":1.0,"reasoning_tokens":895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:17:06.997770+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the two-trial simulation with a binary endpoint at the small end of the paper's grid, such as $p^{(c)}=0.8$, $d=0.2$, and $N^{(s)}=118$; solve $f_k$ from the normal formula (13), and compare the empirical conditional consistency probability over many replications to the nominal 80%. An error larger than simulation noise would show the normal approximation fails there. Also simulate with truly different regional effects, for instance region 1 having effect 0.5 while others have 1.0, and check whether the solved fractions still deliver the claimed probability, which would expose the fixed-effects assumption.","supporting_citations":[{"cited_title":"Ministry of health, habour and helfare of japan basic principles on global clinical trials","cited_arxiv_id":null,"evidence_quote":"Supplies the two consistency criteria and the regional-fraction framework that this paper extends to two pooled trials."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the single-trial normal approximation and Japanese-patient sample-size method that Proposition 4 generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the fixed-effects decision rules and regional sample-size planning that underpin Section 2.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the exact binomial evaluation of the all-regions-positive criterion that Propositions 3 and 6 adopt."},{"cited_title":"Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry","cited_arxiv_id":null,"evidence_quote":"Sets the regulatory requirement of two adequate and well-controlled trials that motivates the two-MRCT design."}],"review_version":1}