Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Multivariate Adjustments for Average Equivalence Testing

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that the conservative multivariate TOST for average equivalence can be made exactly size $\alpha$ in finite samples by raising the test level to a corrected $\alpha^*$ that solves a fixed-point equation, yielding a…

desk verdict Multivariate α-TOST is a real improvement over multivariate TOST for bioequivalence, but the size guarantee rests on an unspecified global optimization of the worst-case null vector. read the letter →

arxiv 2411.16429 v1 pith:3RVVTQZQ submitted 2024-11-25 stat.ME

classification stat.ME MSC 62F0362H1562P10
keywords multivariateequivalencetestingtwoone-sidedtests(TOST)finite-sampleadjustmentsignificancelevelcorrectionbioequivalencetestsizetypeIerrorcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the standard multivariate TOST for average equivalence — checking, outcome by outcome, whether each $100(1-2\alpha)\%$ confidence interval for a mean difference falls inside pre-specified equivalence bounds — is needlessly conservative, and that a corrected significance level $\alpha^*$ restores the test to exactly its nominal size (its intended maximum type-I error rate) at finite samples. The correction is defined as the level at which the worst-case rejection probability over the null space equals $\alpha$, which makes the adjustment depend on the arbitrary correlation between outcomes and on the largest outcome variance. With the true covariance the adjusted test is exactly size $\alpha$; with an estimated covariance the size differs from $\alpha$ by a term of order $o_p(\nu^{-1})$, smaller than the estimation error of the means and variances themselves. Because the correction only ever raises the level, the $\alpha$-TOST's rejection region contains the conventional TOST's, making it uniformly more powerful; simulations across small samples, heterogeneous variances, and correlation structures confirm the gains, and a re-analysis of a ticlopidine hydrochloride bioequivalence study flips a non-equivalence verdict to equivalence. If the claims hold, practitioners get a drop-in finite-sample adjustment that recovers power in multivariate bioequivalence assessment without abandoning the per-outcome interval-inclusion logic.

What carries the argument

The load-bearing object is the size map of the multivariate TOST, $p(\alpha, \theta, \Sigma, \nu, c) = \Pr\left(\bigcap_{j=1}^m \{|\hat\theta_j| \leq c - t_{\alpha,\nu} \hat\sigma_j\}\right)$, computed by integrating a multivariate normal density over outcome-dependent limits against a Wishart density for the covariance estimate (equation 7). Its supremum over the null space $\theta \notin \Theta_1$ defines the test size, and the argument achieving that supremum, $\lambda(\alpha, \Sigma)$ in equation (8), is the quantity that makes the multivariate problem hard: unlike in one dimension, $\lambda$ moves with the level $\alpha$ and with the covariance structure, and it generally has no closed form. The adjustment $\alpha^*(\Omega)$ is the fixed point of the level-to-size map defined by equation (12), and the paper computes it with a two-loop iteration: an inner loop applies the first-order update $\alpha_k = \alpha_{k-1} + \alpha - p(\alpha_{k-1}, \lambda, \Omega, \nu, c)$, which converges exponentially once the size curve satisfies $0 < p'(\gamma) < 2$, and an outer loop re-computes $\lambda$ at each new level until the size equals $\alpha$. Existence of the solution is governed by a threshold on the largest variance, which in the independence case reads $\sigma_{\max} < 2c/\Phi^{-1}(\alpha^{1/m} + 1/2)$ and is argued to be easily satisfied in practice.

What would settle it

Run Algorithm 1 on a heteroscedastic $m = 4$ configuration similar to the case study's covariance, but with the $\lambda$-search performed from many random starting points on the null boundary. If different restarts return measurably different $\hat\alpha^*$ values whose associated true suprema of $p(\hat\alpha^*, \theta, \Sigma, \nu, c)$ over $\theta \notin \Theta_1$ (evaluated by dense grid or large Monte Carlo) deviate from $\alpha$ by more than Monte Carlo error, the global-optimization premise fails and the exact-size claim is not delivered by the algorithm as written. A second check: numerically evaluate $0 < p'(\gamma) < 2$ along the iteration path in that setting — a violation means the exponential-convergence proof does not apply and the output level's exactness is unsupported.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the multivariate TOST size — the supremum of the probability of declaring equivalence over the null hypothesis space — is strictly below $\alpha$, decays sharply as the number of outcomes grows, and is driven mainly by the largest outcome variance and the weakest dependence, so that power often collapses to zero. The proposed fix is the multivariate $\alpha$-TOST: replace $\alpha$ by $\alpha^*(\Omega)$, defined in equation (12) as the value in $[\alpha, 0.5)$ at which $p(\gamma, \lambda(\gamma,\Omega), \Omega, \nu, c) = \alpha$, where $\lambda(\gamma,\Omega)$ is the set of null-space parameter values maximizing the rejection probability. For $\Omega = \Sigma$ the paper proves the procedure is exactly size $\alpha$ (existence conditions supplied, most explicitly for the independence case as $\sigma_{\max} < 2c/\Phi^{-1}(\alpha^{1/m} + 1/2)$); for $\Omega = \hat\Sigma$ it proves $\hat\alpha^* = \alpha^* + o_p(\nu^{-1})$, i.e., the estimated adjustment converges faster than the other random quantities in the problem, and simulations show the feasible procedure's empirical size stays at or below $\alpha$. The monotonicity $\alpha^* \geq \alpha$ then implies the rejection region of the $\alpha$-TOST contains that of the multivariate TOST, so the test is uniformly more powerful, a property the paper exhibits both in an extensive simulation study and in a re-analysis of the ticlopidine hydrochloride bioequivalence data, where the conventional test fails to declare equivalence and the corrected test succeeds.

Load-bearing premise

Everything rests on reliably finding the global worst-case parameter $\lambda$ that maximizes the rejection probability over the null boundary — the algorithm invokes this search without a guarantee of global optimality — together with the 'usually satisfied' condition $0 < p'(\gamma) < 2$ that makes the fixed-point iteration converge to the unique $\alpha^*$; neither is verified for the $m = 4$ settings of the case study.

Editorial extensions

If this is right

  • A multivariate bioequivalence assessment (e.g., $C_{\max}$ and AUC jointly) can use the $\alpha$-TOST at level $\alpha^*$ instead of $\alpha$, recovering power that the per-outcome interval check loses, without changing the interpretation of the intervals.
  • The method is finite-sample: at the population level it is exact for any $\nu$, and the estimated version's size error is negligible relative to the estimation error of means and variances.
  • Settings with large variances on any outcome — the highly-variable-drug scenario — gain the most, since the size gap that $\alpha^*$ corrects is driven by the largest $\sigma_j$.
  • In the case study the correction is mild ($\hat\alpha^* \approx 0.058$) yet sufficient: the $\alpha$-TOST declares bioequivalence of the two ticlopidine formulations where the conventional TOST cannot.
  • Because $\alpha^* \geq \alpha$ always, the $\alpha$-TOST is a uniform power improvement, so no scenario exists in which the corrected test is harder to pass than the conventional one (at the population level).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same fixed-point scheme transfers to asymmetric equivalence margins or to other intersection-union tests, as long as the null-space argmax $\lambda$ can still be evaluated; the paper only develops the symmetric-margin case.
  • The practical guarantee depends on the $\lambda$-search being global; a cheap robustness check for applied users is to restart the optimization from several boundary points and confirm $\hat\alpha^*$ is stable, a diagnostic the paper's implementation does not report.
  • The paper's conjecture that the feasible procedure errs on the conservative side (empirical size below $\alpha$) is itself a regulatory asset: for approval decisions, a size-controlled test that is slightly conservative is safer than one that is exact but numerically fragile.
  • Replacing $\hat\Sigma$ by a robust covariance estimate in the $\alpha^*$ map is a natural follow-up, since the paper flags outliers as the main open threat to the size guarantees.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a multivariate extension of the univariate α-TOST procedure for average equivalence testing. For a given covariance matrix Ω, the adjusted level α*(Ω) is defined in Eq. (12) as the zero of the size function p(γ, λ(γ, Ω), Ω, ν, c) − α, where λ(γ, Ω) is the supremum point of the rejection probability over the null parameter space. The authors claim that, when Ω is the true covariance, the resulting procedure is exactly size α, that the estimated version satisfies pα* = α* + op(ν^{−1}), and that, because α*(Ω) ≥ α, the multivariate α-TOST is uniformly more powerful than the conventional multivariate TOST. The paper presents an iterative algorithm, a simulation study with m ∈ {2,4}, and a re-analysis of the ticlopidine hydrochloride bioequivalence case study with m = 4.

Significance. If the computational and regularity premises hold, the paper gives practitioners a finite-sample size-corrected multivariate equivalence test with power superior to the conventional TOST, which is known to become very conservative as m grows. The extension is nontrivial because the argmax point λ in Eq. (8) depends on the candidate level γ, requiring a nested optimization. The paper provides a formal asymptotic argument in Appendix D, a convergence argument in Appendix E, and an extensive simulation study, and it makes an R implementation available in the cTOST package. However, the central size guarantee is conditional on the ability to compute the global supremum in Eq. (8), and this step is not specified or verified. The claimed uniform power advantage is rigorously justified only for the population-level adjustment, not for the feasible estimated version.

major comments (4)
  1. [§4.2, Algorithm 1; Eqs. (8)–(12)] The size-α property is only as strong as the computation of λ(γ, Ω) in Eq. (8), which is a non-convex supremum over the null boundary of the m-dimensional hypercube. Algorithm 1 invokes 'compute λ(α*(r))' in steps 1 and 10 (and Algorithm 2 does the same in steps 2 and 14) without specifying a global optimizer or providing evidence that the routine used attains the supremum. Moreover, each evaluation of p in Eq. (7) is a high-dimensional integral over a Wishart distribution, typically approximated by Monte Carlo, so the fixed-point iteration is solving a random equation unless the Monte Carlo error is controlled. If the optimization returns a local maximizer λ_loc with p(γ, λ_loc) < sup_{θ ∉ Θ₁} p(γ, θ), then the zero of Eq. (12) is larger than the true size-α level, and the resulting test is liberal at the actual worst-case null parameter. The paper should either specify a certified global optimization method, provide a deterministic or error-bounded procedure for Eq. (7), or include a numerical verification for the m = 4 case study and the simulation scenarios.
  2. [Appendix E, condition 0 < p'(γ) < 2] The convergence and uniqueness proof for the inner fixed-point iteration relies on the assumption that 0 < p'(γ) < 2 on the solution space A, which is stated as 'usually satisfied' without verification. This condition is not checked for the m = 4 case study or for any of the simulation settings in Table 1. The existence of α* is also addressed only through a sufficient condition derived under independence, Eq. (13), while the correlated and heteroscedastic settings are justified heuristically. Without a verifiable condition or a fallback numerical procedure, the algorithm could fail to converge or converge to a point that is not the unique solution of Eq. (12), which would break the claimed size control.
  3. [§4.1, 'uniformly more powerful' claim] The statement that 'its rejection region cannot be smaller than that of the multivariate TOST, which makes the former uniformly more powerful' is proved for the population-level adjusted level α*(Σ), where α*(Σ) ≥ α follows by construction. For the feasible procedure based on pα*, the paper does not establish that pα* ≥ α almost surely, nor does it provide a finite-sample bound on the difference. Section 5 only reports empirical size below α in selected scenarios and says that this conservative behavior is conjectured to extend to the multivariate framework. The uniform-power claim should be explicitly restricted to the population-level version, or a proof or rigorous bound for the feasible version should be supplied.
  4. [Appendix A, structure of λ] The asymptotic size proof in Appendix A assumes that, when max_{i≠j} |ρ_{ij}| < 1, the vector λ contains m−1 components in (−c, c) and exactly one component equal to c. This structural fact about the argmax in Eq. (8) is asserted without proof for general covariance matrices. It underlies both the asymptotic size calculation and the κλ trajectories used in the simulations. The authors should either prove this characterization or state it as a lemma with the needed regularity conditions on Σ.
minor comments (5)
  1. [Table 1] The 'Results in' row of Table 1 refers to 'Figure 7' for Simulation 1, but the simulation results are displayed in Figure 6, while Figure 7 contains the case-study scatterplots; this cross-reference should be corrected.
  2. [Figures 6, A.1–A.3] The horizontal axis in the simulation plots is labeled 'Index NA' in the displayed version; it should be labeled 'κ' uniformly across panels.
  3. [§4.1, existence discussion] The discussion after Eq. (13) uses the independence-based condition to argue existence of α* in the correlated case study; the logical status of this inference should be clarified, since the condition is sufficient but not necessary in that setting.
  4. [§4.2, computation time] The statement that computing the multivariate α-TOST requires 'between 1 to 10 seconds' would be more useful with details on the machine, the Monte Carlo sample size, and the optimization tolerance used.
  5. [Data availability statement] The data and code are available in a GitHub repository, but a versioned release or a persistent archive would improve reproducibility; the authors should also report the versions of R and the cTOST package used for the results.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the size-α guarantee is a transparent construction (Eq. 12), and the multivariate extension is not a restatement of the authors' univariate α-TOST.

full rationale

The paper's central object, α*(Ω), is defined in Eq. (12) as the level γ at which the size function p(γ, λ(γ,Ω), Ω, ν, c) equals α; the statement that using α* gives a size-α test is therefore a direct consequence of the definition, and the paper says so explicitly ('By definition, when a solution to (12) exists...'). This is a constructive design rather than a fitted quantity presented as a prediction: the analytical content is in the existence/uniqueness conditions (Appendix C/E), the recursive dependence of λ on γ, and the Monte Carlo evaluation of p, none of which is an input-output disguise. The multivariate method extends the authors' univariate α-TOST (Boulaguiem et al.), but the multivariate size function (7), the argmax problem (8), and the nested fixed-point algorithm are new; no multivariate claim reduces to the univariate paper's equations by reparametrization. The op(ν^{-1}) result in Appendix D uses standard convergence rates; the citation to the authors' prior work for (A.11)-(A.12) supplies textbook stochastic-order facts and is not the load-bearing justification of the adjustment. No simulation or case-study value is fed back into the derivation, and the feasible level pα* is a plug-in estimate whose empirical size is evaluated externally. The main open issues—global optimization of the non-convex λ(γ,Ω) in Algorithm 1 and the 'usually satisfied' condition 0 < p'(γ) < 2—are numerical/regularity risks, not circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The method introduces no new physical or statistical entities and fits no free parameters. It relies on the canonical normal-Wishart model, symmetric equal margins, and regularity conditions on the power function. The adjusted level α* is the method's output, not a fitted nuisance parameter.

assumptions (4)
  • domain assumption Canonical normal-Wishart model: θ̂ ~ N_m(θ, Σ), νΣ̂ ~ W_m(ν, Σ), with θ̂ and Σ̂ independent (Eq. (1)).
    Used throughout the paper as the data-generating mechanism for the canonical equivalence problem; it covers parallel and crossover designs but assumes normality and Wishart covariance estimates, which may not hold for discrete outcomes like tmax.
  • domain assumption Symmetric equal equivalence margins: b = -a = c for all outcomes (Section 2).
    The paper restricts the alternative space to a hypercube with a single margin c. Asymmetric or outcome-specific margins are not treated.
  • standard math Regularity conditions in Appendix D for the asymptotic result: uniform convergence of the derivatives of the power function p with respect to α and covariance parameters.
    Used to prove pα* = α* + op(ν^{-1}). Standard but not verified for the specific settings studied.
  • domain assumption Contraction condition 0 < p'(γ) < 2 on the power function (Appendix E).
    Assumed to hold to ensure the fixed-point iteration for α* converges exponentially fast; the paper states it is 'usually satisfied' but provides no verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multivariate Adjustments for Average Equivalence Testing." pith.science (2026). https://pith.science/paper/3RVVTQZQ

@misc{pith2026241116429,
  author       = {Pith},
  title        = {Pith review of: Multivariate Adjustments for Average Equivalence Testing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3RVVTQZQ}},
  note         = {Machine review of arXiv:2411.16429}
}
abstract

Multivariate (average) equivalence testing is widely used to assess whether the means of two conditions of interest are `equivalent' for different outcomes simultaneously. The multivariate Two One-Sided Tests (TOST) procedure is typically used in this context by checking if, outcome by outcome, the marginal $100(1-2\alpha$)\% confidence intervals for the difference in means between the two conditions of interest lie within pre-defined lower and upper equivalence limits. This procedure, known to be conservative in the univariate case, leads to a rapid power loss when the number of outcomes increases, especially when one or more outcome variances are relatively large. In this work, we propose a finite-sample adjustment for this procedure, the multivariate $\alpha$-TOST, that consists in a correction of $\alpha$, the significance level, taking the (arbitrary) dependence between the outcomes of interest into account and making it uniformly more powerful than the conventional multivariate TOST. We present an iterative algorithm allowing to efficiently define $\alpha^{\star}$, the corrected significance level, a task that proves challenging in the multivariate setting due to the inter-relationship between $\alpha^{\star}$ and the sets of values belonging to the null hypothesis space and defining the test size. We study the operating characteristics of the multivariate $\alpha$-TOST both theoretically and via an extensive simulation study considering cases relevant for real-world analyses -- i.e.,~relatively small sample sizes, unknown and heterogeneous variances, and different correlation structures -- and show the superior finite-sample properties of the multivariate $\alpha$-TOST compared to its conventional counterpart. We finally re-visit a case study on ticlopidine hydrochloride and compare both methods when simultaneously assessing bioequivalence for multiple pharmacokinetic parameters.

Figures

Figures reproduced from arXiv: 2411.16429 by the authors.

Figure 1
Figure 1. Target parameter space in two bivariate equivalence scenarios, respectively considering independent and p [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Size of the multivariate TOST (y-axis) as a function of [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Target parameter space for the bivariate TOST (left panel) and bivariate [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Colour-coded probabilities of the bivariate TOST rejecting H [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: 3D-heatmap of the probabilities of rejecting the null hypothesis (z-axis) in a bivariate equivalence case (x [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Probability of rejecting the null hypothesis (y-axis) as a function of [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Scatterplots, kernel density estimates and correlations of the (logarithmically transformed) differences be [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Intervals showing the target parameter 100 [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Global p-Values in Multi-Design Studies

    stat.ME 2025-07 conditional novelty 6.0 of 10

    The g-value aggregates p-values across multiple analysis strategies by rescaling the largest p-value with a null-distribution-derived correction, controlling type I error while preserving power.

Reference graph

Works this paper leans on

53 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [1]

    Testing Statistical Hypothesis

    Lehmann, E.L. Testing Statistical Hypothesis. Wiley, New York, 1959

  2. [2]

    A test of an experimental hypothesis of negligible difference between means

    Bondy, W.H. A test of an experimental hypothesis of negligible difference between means. The American Statistician, 23(5):28–30, 1969

  3. [3]

    Use of confidence intervals in analysis of comparative bioavailability trials

    Westlake, W.J. Use of confidence intervals in analysis of comparative bioavailability trials. Journal of Pharma- ceutical Sciences, 61(8):1340–1341, 1972

  4. [4]

    Bioavailability – a problem in equivalence

    Metzler, C. Bioavailability – a problem in equivalence. Journal of Pharmaceutical Sciences , 30:309–317, 1974

  5. [5]

    Symmetrical confidence intervals for bioequivalence trials

    Westlake, W.J. Symmetrical confidence intervals for bioequivalence trials. Biometrics, 32:741–744, 1976

  6. [6]

    Equivalence tests: A practical primer for t tests, correlations, and meta-analyses

    Lakens, D. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science , 8:355–362, 2017

  7. [7]

    “snap on” or not? a validation on the measurement tool in a virtual reality application

    Sureshkumar, H., Xu, R., Erukulla, N., Wadhwa, A., and Zhao, L. “snap on” or not? a validation on the measurement tool in a virtual reality application. Journal of Digital Imaging , 35:692–703, 2022

  8. [8]

    not different

    O’Brien, M.W. and Kimmerly, D.S. Is “not different” enough to conclude similar cardiovascular responses across sexes? American Journal of Physiology-Heart and Circulatory Physiology , 322:H355–H358, 2022

Show all 53 references
  1. [9]

    Similarities and differences in the neurodevelopmental outcome of children with congenital heart disease and children born very preterm at school entry

    Wehrle, F.M., Bartal, T., Adams, M., Bassler, D., Hagmann, C.F., Kretschmar, O., Natalucci, G., and Latal, B. Similarities and differences in the neurodevelopmental outcome of children with congenital heart disease and children born very preterm at school entry. The Journal of...

  2. [10]

    Compara- tive efficacy of Tapentadol versus Tapentadol plus Duloxetine in patients with chemotherapy-induced peripheral neuropathy

    Sansone, P., Giaccari, L.G., Aurilio, C., Coppolino, F., Passavanti, M.B., Pota, V., and Pace, M.C. Compara- tive efficacy of Tapentadol versus Tapentadol plus Duloxetine in patients with chemotherapy-induced peripheral neuropathy. Cancers, 14:4002, 2022

  3. [11]

    No evidence for motor-recovery-related cortical connectivity changes after stroke using resting-state fmri

    Branscheidt, M., Ejaz, N., Xu, J., Widmer, M., Harran, M.D., Cort´ es, J.C., Kitago, T., Celnik, P., Hernandez- Castillo, C., Diedrichsen, J., Luft, A., and Krakauer, J.W. No evidence for motor-recovery-related cortical connectivity changes after stroke using resting-state fmr...

  4. [12]

    Risk-taking for others: An experiment on the role of moral discussion

    Feri, F., Giannetti, C., and Guarnieri, P. Risk-taking for others: An experiment on the role of moral discussion. Journal of Behavioral and Experimental Finance , 37:100735, 2023

  5. [13]

    A 2 million-person, campaign-wide field experiment shows how digital advertising affects voter turnout

    Aggarwal, M., Allen, J., Coppock, A., Frankowski, D., Messing, S., Zhang, K., Barnes, J., Beasley, A., Hantman, H., and Zheng, S. A 2 million-person, campaign-wide field experiment shows how digital advertising affects voter turnout. Nature Human Behaviour , pages 1–10, 2023

  6. [14]

    Equivalence tests – a review

    Meyners, M. Equivalence tests – a review. Food Quality and Preference, 26:231–245, 2012

  7. [15]

    Equivalence testing for psychological research: A tutorial

    Lakens, D., Scheel, A.M., and Isager, P.M. Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science , 1:259–269, 2018. 20

  8. [16]

    Myths and methodologies: The use of equivalence and non-inferiority tests for interventional studies in exercise physiology and sport science

    Mazzolari, R., Porcelli, S., Bishop, D.J., and Lakens, D. Myths and methodologies: The use of equivalence and non-inferiority tests for interventional studies in exercise physiology and sport science. Experimental Physiology, 107:201–212, 2022

  9. [17]

    Wang, K., Li, Y., Chen, B., Chen, H., Smith, D.E., Sun, D., Feng, M.R., and Amidon, G.L. In vitro predictive dissolution test should be developed and recommended as a bioequivalence standard for the immediate-release solid oral dosage forms of the highly variable mycophenolate...

  10. [18]

    The influence of different evaluation techniques on the results of interlaboratory comparisons

    Linsinger, T.P.J., Kandler, W., Krska, R., and Grasserbauer, M. The influence of different evaluation techniques on the results of interlaboratory comparisons. Accreditation and Quality Assurance, 3:322–327, 1998

  11. [19]

    and Small, D.S

    Fogarty, C.B. and Small, D.S. Equivalence testing for functional data with an application to comparing pulmonary function devices. The Annals of Applied Statistics , pages 2002–2026, 2014

  12. [20]

    Knight, R., Dritsaki, M., Mason, J., Perry, D.C., and Dutton, S.J. The forearm fracture recovery in children evaluation (force) trial: statistical and health economic analysis plan for an equivalence randomized controlled trial of treatment for torus fractures of the distal ra...

  13. [21]

    Concordance rate of a four-quadrant plot for repeated measurements

    Hiraishi, M., Tanioka, K., and Shimokawa, T. Concordance rate of a four-quadrant plot for repeated measurements. BMC Medical Research Methodology, 21:1–16, 2021

  14. [22]

    Improved procedures and computer programs for equivalence assessment of correlation coefficients

    Shieh, G. Improved procedures and computer programs for equivalence assessment of correlation coefficients. PLoS ONE, 16(5):e0252323, 2021

  15. [23]

    Assessing statistical similarity in dietary intakes of women of reproductive age in bangladesh.Maternal & Child Nutrition, 17(2):e13086, 2021

    Wable Grandner, G., Dickin, K., Kanbur, R., Menon, P., Rasmussen, K.M., and Hoddinott, J. Assessing statistical similarity in dietary intakes of women of reproductive age in bangladesh.Maternal & Child Nutrition, 17(2):e13086, 2021

  16. [24]

    Multivariate equivalence testing for food safety assessment

    Leday, G.G., Engel, J., Vossen, J.H., de Vos, R.C., and van der Voet, H. Multivariate equivalence testing for food safety assessment. Food and Chemical Toxicology, 170:113446, 2022

  17. [25]

    Ultrastructural calibration model for proficiency testing

    Aoki, R., Le˜ ao, D., Bustamante, J.P.M., and Vilca, F. Ultrastructural calibration model for proficiency testing. Journal of Applied Statistics , 50(5):1037–1059, 2023

  18. [26]

    An item sorting heuristic to derive equivalent parallel test versions from multivariate items

    G¨ obel, N., Cazzoli, D., Gutbrod, C., M¨ uri, R.M., and Eberhard-Moscicka, A.K. An item sorting heuristic to derive equivalent parallel test versions from multivariate items. PLoS ONE, 18(4):e0284768, 2023

  19. [27]

    Bioequivalence of ticlopidine hydrochloride administered in single dose to healthy volunteers

    Marzo, A., Dal Bo, L., Rusca, A., and Zini, P. Bioequivalence of ticlopidine hydrochloride administered in single dose to healthy volunteers. Pharmacological Research, 46(5):401–407, 2002

  20. [28]

    Equivalence analyses of dissolution profiles with the mahalanobis distance

    Hoffelder, T. Equivalence analyses of dissolution profiles with the mahalanobis distance. Biometrical Journal, 61 (5):1120–1137, 2019. 21

  21. [29]

    New model–based bioequivalence statistical approaches for pharmacokinetic studies with sparse sampling

    Loingeville, F., Bertrand, J., Nguyen, T.T., Sharan, S., Feng, K., Sun, W., Han, J., Grosser, S., Zhao, L., Fang, L., et al. New model–based bioequivalence statistical approaches for pharmacokinetic studies with sparse sampling. The AAPS Journal , 22:1–8, 2020

  22. [30]

    Testing Statistical Hypotheses of Equivalence and Noninferiority

    Wellek, S. Testing Statistical Hypotheses of Equivalence and Noninferiority . CRC Press, 2010

  23. [31]

    A multivariate test for population bioequivalence

    Chervoneva, I., Hyslop, T., and Hauck, W.W. A multivariate test for population bioequivalence. Statistics in Medicine, 26(6):1208–1223, 2007

  24. [32]

    A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability

    Schuirmann, D.J. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics , 15(6):657–680, 1987

  25. [33]

    and Jaki, T

    Pallmann, P. and Jaki, T. Simultaneous confidence regions for multivariate bioequivalence. Statistics in Medicine, 36(29):4585–4603, 2017

  26. [34]

    Optimal confidence sets, bioequivalence, and the limacon of pascal

    Brown, L.D., Casella, G., and Gene Hwang, J. Optimal confidence sets, bioequivalence, and the limacon of pascal. Journal of the American Statistical Association , 90(431):880–889, 1995

  27. [35]

    Optimal confidence sets for testing average bioequivalence

    Tseng, Y.L. Optimal confidence sets for testing average bioequivalence. Test, 11(1):127–141, 2002

  28. [36]

    and Hsu, J.C

    Berger, R.L. and Hsu, J.C. Bioequivalence trials, intersection-union tests and equivalence confidence sets. Statis- tical Science, 11:283–319, 1996

  29. [37]

    An unbiased test for the bioequivalence problem

    Brown, L.D., Hwang, J.G., and Munk, A. An unbiased test for the bioequivalence problem. The Annals of Statistics, pages 2345–2367, 1997

  30. [38]

    Finite sample corrections for average equivalence testing

    Boulaguiem, Y., Quartier, J., Lapteva, M., Kalia, Y.N., Victoria-Feser, M.P., Guerrier, S., and Couturier, D.L. Finite sample corrections for average equivalence testing. Statistics in Medicine , 43(5):833–854, 2024

  31. [39]

    Statistical tests for multivariate bioequivalence

    Wang, W., Gene Hwang, J., and Dasgupta, A. Statistical tests for multivariate bioequivalence. Biometrika, 86 (2):395–402, 1999

  32. [40]

    and Tothfalusi, L

    Endrenyi, L. and Tothfalusi, L. Bioequivalence for highly variable drugs: regulatory agreements, disagreements, and harmonization. Journal of Pharmacokinetics and Pharmacodynamics , 46:117–126, 2019

  33. [41]

    Multiparameter hypothesis testing and acceptance sampling

    Berger, R.L. Multiparameter hypothesis testing and acceptance sampling. Technometrics, 24(4):295–300, 1982

  34. [42]

    Power of the two one-sided tests procedure in bioequivalence

    Phillips, K. Power of the two one-sided tests procedure in bioequivalence. Journal of Pharmacokinetics and Biopharmaceutics, 18:137–144, 1990

  35. [43]

    A special case of a bivariate non-central t-distribution

    Owen, D.B. A special case of a bivariate non-central t-distribution. Biometrika, 52:437–446, 1965

  36. [44]

    Power for testing multiple instances of the two one-sided tests procedure

    Phillips, K.F. Power for testing multiple instances of the two one-sided tests procedure. The International Journal of Biostatistics , 5(1), 2009

  37. [45]

    Testing Statistical Hypothesis, 2nd ed

    Lehmann, E.L. Testing Statistical Hypothesis, 2nd ed . Wiley, New York, 1986. 22

  38. [46]

    and Zhou, X.H

    Deng, Y. and Zhou, X.H. Methods to control the empirical type I error rate in average bioequivalence tests for highly variable drugs. Statistical Methods in Medical Research, 29:1650–1667, 2020

  39. [47]

    Guidance for industry, statistical approaches to establishing bioequivalence

    Food and Drugs Administration. Guidance for industry, statistical approaches to establishing bioequivalence. http://www.fda.gov/cder/guidance/index.htm, 2001

  40. [48]

    Guideline on the investigation of bioequivalence-cpmp

    European Medicine Agency. Guideline on the investigation of bioequivalence-cpmp. Technical report, EWP/QWP/1401/98 Rev. 1, 2010

  41. [49]

    and Sch¨ utz, H

    Labes, D. and Sch¨ utz, H. Inflation of type i error in the evaluation of scaled average bioequivalence, and a method for its control. Pharmaceutical Research, 33:2805–2814, 2016

  42. [50]

    GGally: Extension to ’ggplot2’ , 2024

    Schloerke, B., Cook, D., Larmarange, J., Briatte, F., Marbach, M., Thoen, E., Elberg, A., and Crowley, J. GGally: Extension to ’ggplot2’ , 2024. URL https://CRAN.R-project.org/package=GGally. R package version 2.2.1

  43. [51]

    Asymptotic Statistics

    van der Vaart, A.W. Asymptotic Statistics . Cambridge University Press, Cambridge, 2000

  44. [52]

    Principles of Mathematical Analysis

    Rudin, W. Principles of Mathematical Analysis . McGraw-Hill, 2nd edition, 1976

  45. [53]

    α. Notice that, given α, λ (as defined in Equation 8), Σ, ν and c, the size can be written as follows ppα, λ, Σ, ν,cq“ Pr ˜ mč j“1 ! |pθj|ă c´ tα,νpσj )¸ “ Pr ˜ mč j“1

    Federer, H. Geometric Measure Theory. Springer, 2014. 23 Appendix A. Asymptotic Size of the Multivariate TOST Here we demonstrate that the multivariate TOST is size- α asymptotically, namely that lim νÑ8 ppα, λ, Σ, ν,cq“ α. Notice that, given α, λ (as defined in Equation 8), Σ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.