Pith. sign in

REVIEW 3 major objections 6 minor 19 references

Regional consistency evaluation and sample size calculation under two MRCTs

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper derives double-integral formulas for regional consistency after two pivotal MRCTs are both significant, and shows these formulas can be solved to set regional sample fractions.

desk verdict Genuine extension of MHLW regional-consistency methods to two-pivotal-MRCT planning, with solid simulations, but the central pooling identity is an unstated approximation that needs fixing. read the letter →

arxiv 2411.15567 v2 pith:E3MLFB2L submitted 2024-11-23 stat.AP

classification stat.AP MSC 62P1062K05
keywords multi-regionalclinicaltrialsregionalconsistencyevaluationsamplesizecalculationfixedeffectsmodelMHLWcriteriaprobabilitypooledanalysisbinaryendpoints
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Drug registration often requires two pivotal multi-regional trials, and a region may need to join both to reach a required subpopulation. This paper asks how a region should split its enrollment across the two trials so that, once both trials show an overall effect, the region's pooled estimate credibly matches the pooled overall estimate. Under the fixed-effects model with a single common treatment effect, it derives explicit approximations for the probability of regional consistency, expressed as double integrals in the two standardized trial statistics. Solving those integrals gives the regional sample fractions in each trial that attain a desired consistency probability, and the paper's simulations show the empirical probability lands within about one percentage point of the target. The practical payoff is a design tool, implemented in an R package, for regions that must join both pivotal MRCTs.

What carries the argument

The machinery is the joint asymptotic normality of the pooled regional and overall estimators, combined with conditioning on both study test statistics exceeding $z_{1-\alpha}$. The paper forms $D_{k,\mathrm{pool}}=w^{(1)}D_k^{(1)}+w^{(2)}D_k^{(2)}$ and $D_{\mathrm{pool}}=w^{(1)}D^{(1)}+w^{(2)}D^{(2)}$, then uses the identity $\sigma_d^{2(s)} = d^2_{(s)}/(z_{1-\alpha}+z_{1-\beta_s})^2$ to turn each rejection event $T^{(s)} > z_{1-\alpha}$ into $Z^{(s)} > -z_{1-\beta_s}$. That converts the conditional consistency probability into a double integral with no unknown parameters beyond design inputs, and the required regional fractions depend on $\alpha$, the powers, effect sizes, variances, and randomization ratios, but not on the total sample sizes $N^{(s)}$.

What would settle it

Run the two-trial simulation with a binary endpoint at the small end of the paper's grid, such as $p^{(c)}=0.8$, $d=0.2$, and $N^{(s)}=118$; solve $f_k$ from the normal formula (13), and compare the empirical conditional consistency probability over many replications to the nominal 80%. An error larger than simulation noise would show the normal approximation fails there. Also simulate with truly different regional effects, for instance region 1 having effect 0.5 while others have 1.0, and check whether the solved fractions still deliver the claimed probability, which would expose the fixed-effects assumption.

Watch

Extended reading notes

Core claim

The central result is Proposition 4. For two independent pivotal trials with treatment differences $d^{(s)}>0$, the authors approximate the conditional consistency probability $$\Pr\!\left(D_{k,\mathrm{pool}} \ge \pi D_{\mathrm{pool}} \mid $T^{{(1)}}$ > z_{1-\$\alpha$},\, $T^{{(2)}}$ > z_{1-\$\alpha$}\right)$$ by a double integral over the standardized trial statistics $u,v$, whose integrand is a normal CDF and whose normalizing denominator is $(1-\beta_1)(1-\beta_2)$, as in equation (13). Because this probability is monotone in the combined fraction quantity $\zeta$, a planner can solve for regional fractions $f_k^{(1)}$ and $f_k^{(2)}$ to hit a target consistency probability, and the combined regional sample size is minimized when the fraction ratio obeys equation (14). The same conditioning device is applied to the second MHLW criterion, where all regions must show the same trend, yielding Proposition 5, and exact binomial versions are given for binary endpoints in Propositions 3 and 6. The paper reports that with these solved fractions the empirical consistency probability tracks the nominal 80% to within roughly one percent across continuous and binary response settings.

Load-bearing premise

The results assume a single true treatment effect shared by all regions and both trials, and that the test statistics are close enough to normal for the double integral to hold.

Editorial extensions

If this is right

  • If formula (13) is right, a region enrolling in both trials no longer needs to carry a large share of either trial; in the paper's worked example, an 80% consistency probability is met with only 10.9% of each study's sample under equal fractions.
  • Because many fraction pairs $(f_k^{(1)}, f_k^{(2)})$ achieve the same consistency probability, a planner can trade enrollment between studies, such as (8%, 17.4%) or (9%, 14.1%) in the example, while the minimum total regional size pins down the ratio in (14).
  • The pooled criterion is substantially more efficient than checking each trial separately: in Remark 9, the required fraction drops from 46.6% per trial to 15.4% once the two trials are pooled and both are significant.
  • For binary endpoints, the normal double-integral approximation under-sizes the region, and the exact binomial formulas are needed; Example 4 shows $f_1 = 6.0\%$ rather than the normal-approximation value of 4.4%.
  • The methods come with an R package, so the fraction solving and consistency-probability evaluation are directly usable in practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The conditioning device generalizes naturally: if more than two pivotal trials must all be significant, the same construction gives a higher-dimensional integral, and the equal-fraction optimum under homogeneity should continue to hold, though this is not stated in the paper.
  • The binary-endpoint correction suggests that the normal approximation should be stress-tested for time-to-event endpoints, where test statistics are not normal in moderate samples; the paper flags this as future work.
  • A sponsor facing enrollment constraints could replace the minimum-total-sample objective with other constraints, such as a cap on one trial's fraction, because every consistency-probability level has a curve of feasible fraction pairs even though the paper only solves the minimum-size objective.
  • The paper's own caution implies a boundary: if region-specific covariate distributions shift differently in the two studies, the pooled estimate is not a valid consistency target, and an extension could quantify the bias by modeling the shift explicitly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper addresses regional consistency evaluation and sample size calculation for two multi-regional clinical trials (MRCTs). The authors first review existing criteria for a single MRCT under a fixed-effects model and present unified approximations for the MHLW type I and type II consistency probabilities (Propositions 1–3). They then extend these criteria to the setting where a region participates in two pivotal MRCTs, proposing pooled-data criteria (12) and (16). Proposition 4 gives an approximate conditional consistency probability for criterion (12) as a double integral; Propositions 5 and 6 give the corresponding results for the all-regions criterion and for binary endpoints. Regional sample size fractions are solved numerically, with a closed-form optimality condition (14) for minimizing the combined regional sample size. Simulations with B = 100,000 replications and a hypothetical example illustrate the methods, and an R package is provided.

Significance. The two-MRCT extension fills a genuine practical gap, as regions sometimes need to enroll across both pivotal trials. The paper provides a clear framework, ready-to-use numerical integration schemes, extensive simulations, and an R package; the one-MRCT results are standard and the empirical CP values are close to nominal in the reported grids. The key novelty, however, rests on Proposition 4, and that derivation contains an algebraic misspecification: the identities in Remark 9 hold only under equal randomization ratios and equal regional fractions across studies. For the general case (e.g., Tables 5–6 with r(1)=1, r(2)=2), Eq. (13) does not evaluate the consistency probability of the pooled estimator defined in Eq. (11). This must be corrected before the method can be recommended for general use.

major comments (3)
  1. [2.3, Remark 9 and Appendix A.3 (Eq. 13)] The identities D_{k,pool}=w^{(1)}D_k^{(1)}+w^{(2)}D_k^{(2)} and D_{pool}=w^{(1)}D^{(1)}+w^{(2)}D^{(2)} stated in Remark 9 and used in the proof of Proposition 4 are not identities for the estimators defined in Eq. (11). From (11), D_{pool} is a weighted average with treatment-arm weights N_t^{(s)}/(N_t^{(1)}+N_t^{(2)}) and control-arm weights N_c^{(s)}/(N_c^{(1)}+N_c^{(2)}); these equal w^{(s)}=N^{(s)}/(N^{(1)}+N^{(2)}) only if r^{(1)}=r^{(2)}. Similarly, D_{k,pool} has weights proportional to f_k^{(s)}N_t^{(s)} and f_k^{(s)}N_c^{(s)}, which equal w^{(s)} only if, additionally, f_k^{(1)}=f_k^{(2)}. Consequently, the conditional mean (1−π)(w^{(1)}D^{(1)}+w^{(2)}D^{(2)}) and the variance terms ((f_k^{(1)})^{-1}−1)(w^{(1)}σ_d^{(1)})^2 + ((f_k^{(2)})^{-1}−1)(w^{(2)}σ_d^{(2)})^2 in Eq. (13) are not the moments of D_{k,pool}−πD_{pool} under the definition (11). Tables 5–6 and S3–S4 use exactly the unequal-ratio case r^{(1)}=1, r^{(2)}=2 (and unequal f_k^{(s)}), so the close agreement of empirical CP does not validate the derivation; it only indicates that the misspecification is numerically small in those scenarios. The authors should either restrict Propositions 4–6 to settings where the weighting identities hold (r^{(1)}=r^{(2)} and f_k^{(1)}=f_k^{(2)}) or re-derive the consistency probability and sample-size formulas using the arm-specific weights in (11).
  2. [2.3, Eq. (14) and Remark 6] The minimization condition (14) is a consequence of the variance expression in (13). Because that variance is not the variance of the actual pooled estimator when the weights differ, the reported 'combined sample size minimized' solutions in Tables 5–6 and S3–S4 are not necessarily optimal for the estimator in (11). The authors should re-derive the optimality condition from the correct conditional variance, or state explicitly that (14) optimizes the approximate variance in the w-weighted model rather than the actual pooled estimator.
  3. [2.3, Proposition 6 (Eq. 18)] The event in (18) compares w^{(1)}p̂_{t,k}^{(1)}+w^{(2)}p̂_{t,k}^{(2)} with w^{(1)}p̂_{c,k}^{(1)}+w^{(2)}p̂_{c,k}^{(2)}, which is the w-weighted combination of study-specific regional differences. This is not the same as D_{k,pool} ≥ 0 defined by the pooled estimator in (11) unless the weights coincide as in the first comment. The same correction is needed for the binary-response exact calculation; otherwise Proposition 6 evaluates a different event from the LHS of (16).
minor comments (6)
  1. [Throughout] The acronym 'MHLW' is inconsistently spaced as 'MHL W' in the abstract, introduction, and elsewhere; please make it consistent.
  2. [Example 2] There is a typo: 'respetively' should be 'respectively'.
  3. [Example 4] The line 'p(c,1) = p(c,1) = 0.8' should read 'p(c,1) = p(c,2) = 0.8'.
  4. [Proposition 6] In the sentence defining the binomial distributions, 'b(s)_k ∼ Bin(f(s)_k N(c,s), p(t,s))' should use p(c,s) rather than p(t,s).
  5. [Remark 2] The claim that the approximation error in (6) is O(n^{-1/2}) is not demonstrated in the Appendix; the proof only shows the asymptotic approximation, not the stated rate. Either provide a proof of the rate or soften the claim to 'the approximation is asymptotically justified'.
  6. [Section 4] The sentence 'The global enrollment of Study 1 and Study 2 is sequential' should be 'are sequential'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the consistency-probability formulas are derived from stated distributional assumptions and independently verified by simulation.

full rationale

The paper derives its sample-size and consistency-probability (CP) formulas, especially Equation (13) for two MRCTs, from explicit normal-approximation arguments built on the fixed-effects model, the assumed overall treatment effect, and the sample-size identity for the two-sample test. No parameters are fitted to the data whose probability the formulas then 'predict'; the inputs (alpha, beta, d, sigma, r, f) are design assumptions, and the reported empirical CP values come from independent simulations, not from plugging the formula into itself. The authors' reliance on prior work (MHLW guidance, Ikeda and Bretz, Ko et al., Homma) is either background or an imported, independently published result (e.g., Proposition 3), and the only self-citation (Adall and Xu 2021) is a passing 'See also' reference that does not carry the derivation. The Discussion explicitly flags when the pooled approach is not appropriate (covariate-distributional shifts differing across studies), which is an applicability limitation rather than a circularity. A reviewer concern that the pooling identity in Remark 9 is exact only when randomization ratios are equal is an algebraic-approximation issue about the regime in which Equation (13) is applied; it does not make the derivation equivalent to its own inputs, so it lies outside the circularity score. Overall the derivation chain is self-contained, with assumptions stated rather than smuggled in, and the claimed operating characteristics are checked externally by simulation.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters: the paper is a design-stage methodology; all inputs (d, sigma, r, p(c), alpha, beta, pi, gamma) are assumed design values or MHLW/FDA conventions, not fitted to data. The axioms are the fixed-effects model, large-sample normality, the design-stage identity (20), independence of the two trials, equal randomization ratio across regions, and the Remark 9 decomposition used in the proofs.

assumptions (6)
  • standard math Large-sample normality of D and D_k (equations (3)-(4))
    Invoked throughout Section 2; justified by the central limit theorem for sums of independent samples.
  • domain assumption The fixed effects model: common treatment effect d and common variance across all regions and both studies
    Stated in Section 2.1 and used in all propositions; the Discussion cautions it fails under covariate-shift heterogeneity.
  • domain assumption Design-stage effect identity: sigma_d^2 = d^2 / (z_{1-alpha} + z_{1-beta})^2 (equation (20))
    Follows from the sample size formula (1) and assumes the true effect equals the design effect d; used to convert significance events into Z-threshold events.
  • domain assumption Independence of the two MRCTs
    Stated in Section 2.3; needed for the denominator (1-beta1)(1-beta2) and the product form of the joint density.
  • domain assumption Equal randomization ratio across regions within each study
    Assumed in Section 2.1 and used to justify N_k^{(t)} / N_k^{(c)} = r.
  • ad hoc to paper D_{k,pool} = w^{(1)}D_k^{(1)} + w^{(2)}D_k^{(2)} (Remark 9)
    Invoked in proofs of Propositions 4 and 5; holds exactly only when r^{(1)} = r^{(2)}, but the paper uses it without this qualification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Regional consistency evaluation and sample size calculation under two MRCTs." pith.science (2026). https://pith.science/paper/E3MLFB2L

@misc{pith2026241115567,
  author       = {Pith},
  title        = {Pith review of: Regional consistency evaluation and sample size calculation under two MRCTs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E3MLFB2L}},
  note         = {Machine review of arXiv:2411.15567}
}
read the original abstract

Multi-regional clinical trial (MRCT) has been common practice for drug development and global registration. The FDA guidance `Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry' (FDA, 2019) requires that substantial evidence of effectiveness of a drug/biologic product to be demonstrated for market approval. In the situations where two pivotal MRCTs are needed to establish effectiveness of a specific indication for a drug or biological product, a systematic approach of consistency evaluation for regional effect is crucial. In this paper, we first present some existing regional consistency evaluations in a unified way that facilitates regional sample size calculation under the simple fixed effects model. Second, we extend the two commonly used consistency assessment criteria of MHLW (2007) in the context of two MRCTs and provide their evaluation and regional sample size calculation. Numerical studies demonstrate the proposed regional sample size attains the desired probability of showing regional consistency. A hypothetical example is presented to illustrate the application. We provide an R package for implementation.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 19 canonical work pages

  1. [1]

    Adall, S. W. and J. Xu (2021). Bayes shrinkage estimator for consistency assessment of treatment effects in multi-regional clinical trials. Pharmaceutical Statistics\/ 20\/ (6), 1074--1087

  2. [2]

    Chen, J., H. Quan, B. Binkowitz, S. P. Ouyang, Y. Tanaka, G. Li, S. Menjoge, and E. Ibia (2010). Assessing consistent treatment effect in a multi-regional clinical trial: a systematic review. Pharmaceutical Statistics\/ 9\/ (3), 242--253

  3. [3]

    Chen, X., N. Lu, R. Nair, Y. Xu, C. Kang, Q. Huang, N. Li, and H. Chen (2012). Decision Rules and Associated Sample Size Planning for Regional Approval Utilizing Multiregional Clinical Trials . Journal of Biopharmaceutical Statistics\/ 22\/ (5), 1001--1018

  4. [4]

    Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry

    FDA (2019). Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry

  5. [5]

    Haber, M. (1986). An exact unconditional test for the 2 2 comparative trial. Psychological Bulletin\/ 99\/ (1), 129--132

  6. [6]

    Homma, G. (2023). Cautionary note on regional consistency evaluation in multiregional clinical trials with binary outcomes. Pharmaceutical Statistics\/

  7. [7]

    International Council for Harmonisation E17 General Principles for Planning and Design of Multi-Regional Clinical Trials

    ICH E17 (2017). International Council for Harmonisation E17 General Principles for Planning and Design of Multi-Regional Clinical Trials

  8. [8]

    Ikeda, K. and F. Bretz (2010). Sample size and proportion of Japanese patients in multi-regional trials. Pharmaceutical Statistics\/ 9\/ (3), 207--216

Show all 19 references
  1. [9]

    Tsou, J.-P

    Ko, F.-S., H.-H. Tsou, J.-P. Liu, and C.-F. Hsiao (2010). Sample size determination for a specific region in a multiregional trial. Journal of Biopharmaceutical Statistics\/ 20\/ (4), 870--885

  2. [10]

    Quan, and Y

    Li, G., H. Quan, and Y. Wang (2024). Regional consistency assessment in multiregional clinical trials. Journal of Biopharmaceutical Statistics\/ , 1--13

  3. [11]

    Ministry of health, habour and helfare of japan basic principles on global clinical trials

    MHLW (2007). Ministry of health, habour and helfare of japan basic principles on global clinical trials

  4. [12]

    Quan, H., M. Li, W. J. Shih, S. P. Ouyang, J. Chen, J. Zhang, and P. Zhao (2013). Empirical shrinkage estimator for consistency assessment of treatment effects in multi‐regional clinical trials. Statistics in Medicine\/ 32\/ (10), 1691--1706

  5. [13]

    Quan, H., P. Zhao, J. Zhang, M. Roessner, and K. Aizawa (2010). Sample size considerations for Japanese patients in a multi‐regional trial based on MHLW guidance. Pharmaceutical Statistics\/ 9\/ (2), 100--112

  6. [14]

    Suissa, S. and J. J. Shuster (1985). Exact Unconditional Sample Sizes for the 2 2 Binomial Trial . Journal of the Royal Statistical Society. Series A (General)\/ 148\/ (4), 317

  7. [15]

    Tanaka, Y., G. Li, Y. Wang, and J. Chen (2012). Qualitative Consistency of Treatment Effects in Multiregional Clinical Trials . Journal of Biopharmaceutical Statistics\/ 22\/ (5), 988--1000

  8. [16]

    Chen, and M

    Teng, Z., Y.-F. Chen, and M. Chang (2017). Unified additional requirement in consideration of regional approval for multiregional clinical trials. Journal of Biopharmaceutical Statistics\/ 27\/ (6), 903--917

  9. [17]

    Chang, X

    Tsong, Y., W.-J. Chang, X. Dong, and H.-H. Tsou (2012). Assessment of Regional Treatment Effect in a Multiregional Clinical Trial . Journal of Biopharmaceutical Statistics\/ 22\/ (5), 1019--1036

  10. [18]

    Chien, J.-p

    Tsou, H.-H., T.-Y. Chien, J.-p. Liu, and C.-F. Hsiao (2011). A consistency approach to evaluation of bridging studies and multi-regional trials. Statistics in Medicine\/ 30\/ (17), 2171--2186

  11. [19]

    Xu, X.-J

    Wu, S.-C., J.-F. Xu, X.-J. Zhang, Z.-W. Li, and J. He (2020). Regional consistency and sample size considerations in a multiregional equivalence trial. Pharmaceutical Statistics\/ 19\/ (6), 897--908

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.