Pith. sign in

REVIEW 4 major objections 6 minor 5 references

The 'expected f2' method for comparing dissolution profiles is, according to this position paper, an unjustified variance-penalized statistic that can reject equivalence when the true profiles are identical and should not appear in regulato

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 20:02 UTC pith:TGAPVXFY

load-bearing objection A solid, checkable critique of the 'expected f2' statistic: the bias argument is correct under the standard population-f2 target, and the source-tracing is genuinely useful, but the paper needs a reproducible simulation and should defend why the regulatory estimand is population f2. the 4 major comments →

arxiv 2607.16759 v1 pith:TGAPVXFY submitted 2026-07-18 stat.ME stat.AP

EFSPI CMCSNE SIG position on the 'Expected f2'

classification stat.ME stat.AP MSC 62F0362F4062P10
keywords similarity factor f2expected f2dissolution profile comparisonbias correctionequivalence testingregulatory guidancevariance termbootstrap confidence interval
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the method called 'expected f2' — a statistic for comparing drug dissolution profiles when measurements are highly variable — is not a legitimate statistical procedure. The core technical claim is that the formula inserts the estimated variance term h/n inside the logarithm with a plus sign, where the bias-corrected formulation from which it descends inserts it with a minus sign; since the logarithm is multiplied by −25, this lowers the computed value and makes the statistic stricter as variability grows. The paper further claims the formula has no traceable origin in the references cited by its proponents, that no mathematical derivation accompanies it, and that a bracket ambiguity in the original publication makes the intended formula unclear — an ambiguity that has carried into regulatory guidance. Because a survey of the working group found no scientific advantage and the statistic can reject equivalence when true profiles are identical, the paper concludes the method should not be recommended and should be removed from guidance.

Core claim

The paper's central claim is that 'expected f2' is the mirror image of a bias correction, not an expectation. Where the original bias-corrected estimator computes 100 − 25·log10(1 + g(μ̂) − h(σ̂²)/n), the expected-f2 formula computes 100 − 25·log10(1 + g(μ̂) + h(σ̂²)/n). The negative coefficient on the log means the added variance term always pulls the statistic downward below the conventional f2 point estimate; under high variability the h/n term dominates, so the statistic's acceptance probability is bounded below 1 and can approach 0 even when the true mean difference is zero. The paper supports this with a simulated example where reference and test profiles are identical, conventional f2

What carries the argument

The mechanism is a sign flip inside the defining identity of the statistic. Using the auxiliary functions g(μ̂) = (1/p)Σ(X̄_Ti − X̄_Ri)² and h(σ̂²) = (1/p)Σ(S²_Ti + S²_Ri), the paper contrasts the bias-corrected estimator f̂2,bc = 100 − 25 log10(1 + g − h/n) with the expected-f2 estimator f̂2,exp = 100 − 25 log10(1 + g + h/n). The intended estimand is population f2 = 100 − 25 log10(1 + QMD²), with equivalence defined by f2 > 50, equivalent to quadratic-mean difference QMD < about 10%. Because f2 is a decreasing function of its log argument, inserting h/n makes f̂2,exp ≤ f̂2 and converts a variance correction into a variance penalty; the decision threshold remains fixed at 50, so all of the p

Load-bearing premise

Everything hinges on expected f2 being meant to estimate the population parameter f2 = 100 − 25 log10(1 + QMD²), not to serve as a deliberately conservative regulatory statistic; if that framing is granted, the variance term is an error, but the paper asserts rather than derives this target.

What would settle it

Simulate identical reference and test profiles (true mean difference zero) across a grid of variances and sample sizes; compute the two published readings of expected f2 and bootstrap acceptance rates. The paper predicts acceptance probability collapses toward zero as variance grows even when profiles are identical; any design where acceptance probability goes to 1 with sample size, or stays high despite high variance, would falsify the dominance claim. A second check is archival: locate a derivation or original proposal of the plus-sign formula in the cited references; the paper claims none e

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the paper is right, the two regulatory guidelines that adopted expected f2 for high-variability dissolution comparisons rest on a formula with no derivation and should be revised or withdrawn.
  • Sponsors using the method for highly variable products will systematically underpower equivalence claims; truly equivalent products can fail the comparison, triggering unnecessary in vivo bioequivalence studies.
  • The bracket ambiguity in the original publication means the method as published defines two different statistics, and the two adopted guideline versions inherit different readings; results can differ by more than 8 points in the paper's example.
  • The equal-sample-size restriction means expected f2 cannot be paired with the bias-corrected and accelerated bootstrap interval, because the jackknife step produces unequal bootstrap samples; this blocks the one standard way to get a confidence interval for f2.
  • The paper's sign-reversal argument also undermines the claimed conservatism advantage: a lower Type I error rate is bought entirely by inflating Type II error, not by any improved estimator property.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • My inference: the sign-flip critique only refutes expected f2 as an estimator of population f2. If a defender re-framed it as a deliberately conservative, variance-penalized statistic for regulatory decisions, the plus sign would be by design; the paper does not address that re-framing, so its conclusion is conditional on the estimand being population f2.
  • My inference: the variance-dominance mechanism predicts a sharp power cliff. A testable extension would map acceptance probability as a function of total variance at QMD=0; the paper's claim implies acceptance probability tends to 0 as variance grows, whereas a merely conservative test should approach 1 for identical profiles as n grows.
  • My inference: the same sign-reversal test applies to any 'variance-inflated' similarity metric. The paper's identity suggests checking whether the proposed conservative statistic targets a parameter of the sampling distribution rather than the population difference; if it does not, adding variance will always trade power without a stated estimand.
  • My inference: beyond statistics, the bracket ambiguity propagating into guidance is a governance lesson: regulatory formulas should be accompanied by derivations and machine-readable specification so that sign and grouping errors are caught before adoption.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This position paper by the EFSPI CMCSNE SIG argues that the 'expected f2' method (f̂2,exp), recommended in EMA Q&A 3.13 (2025) and the MFDS guideline (2023) for dissolution profile comparison under high variability, is scientifically unjustified. The authors contend that the formula for f̂2,exp has no traceable origin in the cited literature, that it adds rather than subtracts a variance term inside the logarithm (thereby increasing downward bias relative to the population f2 defined by the equivalence hypotheses in Eq. [1]), that it can produce near-zero power even when the true profiles are identical, and that the Noce et al. (2020) formulation contains an ambiguous bracket structure that has propagated into regulatory guidance. The paper concludes that f̂2,exp should not be recommended and advocates for engagement with health authorities.

Significance. If the paper's central critique holds, it has immediate practical consequences: f̂2,exp is embedded in two current regulatory guidelines, and the paper's logic would imply that these guidelines recommend a method with anti-conservative bias in the wrong direction and potentially unacceptably low power. The elementary mathematical observation underlying the critique is exact: because h(σ̂²)/n ≥ 0 enters the log argument with a negative coefficient, f̂2,exp ≤ f̂2 always, so if the target is the population parameter f2 defined in Eq. [1], the variance correction is indeed in the wrong direction. The paper also cites an independent FDA assessment (Liu et al., 2024) and includes a reproducible-in-principle numerical example. However, the significance is moderated by two gaps: the paper does not quote the regulatory guidelines to establish that f̂2,exp is intended as a point estimator of population f2 (rather than as a deliberately conservative decision statistic), and the numerical evidence is a single simulation with undisclosed settings.

major comments (4)
  1. [Section 4 and §5, Eq. [3]] The 'wrong direction' bias claim is the load-bearing assertion, and it depends on identifying the estimand of f̂2,exp with the population f2 defined via QMD in Eq. [1]. The paper defines this parameter and the equivalence hypotheses itself, but it does not quote EMA Q&A 3.13 or the MFDS guideline to show that these regulators intend f̂2,exp as a point estimator of population f2. If the guidance instead frames f̂2,exp as a conservative statistical decision rule for high-variability situations, adding h(σ̂²)/n is a design choice, not an error. The paper must either provide the relevant regulatory text demonstrating the intended estimand or explicitly limit the bias critique to the population-f2 estimand and acknowledge the alternative interpretation.
  2. [Section 8, Figure [1]] The numerical illustration that 'the maximum power obtained for both f̂2,exp variants is close to 0' is based on a single simulated dataset with no specification of sample size n, number of time points p, variance magnitudes, or bootstrap settings (including number of replications and the method for constructing the lower confidence limit). Without these details, the reader cannot assess whether the result is generic or an artifact of extreme settings. A systematic power evaluation over a grid of QMD values and variance levels, with full simulation parameters, is needed to support the strong claim in the abstract and Section 7 that the method has power approaching zero even at QMD = 0.
  3. [Section 7, point 1] The statement that the Type I error rate of f̂2,exp is 'of course lower than the T1E rate of other approaches' is not well-defined. The T1E rate depends on the entire decision rule (e.g., whether equivalence is concluded when a lower bootstrap confidence limit exceeds 50, or some other rule), and the comparison set of 'other approaches' is not specified. Without a formal definition of the test procedure for f̂2,exp and a clearly defined comparator test, the T1E claim is not falsifiable. This is secondary to the bias/power argument, but it is used as a supporting property and needs precision.
  4. [Section 6] The notation ambiguity in Noce et al. (2020) is a legitimate practical concern, but the paper elevates it to one of 'four fundamental concerns' on par with the bias and power issues. A missing bracket is a typographical/generalization problem that could be resolved by editorial correction; it does not fundamentally invalidate the method. The paper should reclassify this as a practical guidance issue rather than a fundamental statistical flaw, or explain why the ambiguity has substantive consequences beyond calculation differences.
minor comments (6)
  1. [Section 3] Typo: 'distribution of the the TEST population' should be 'the TEST population'.
  2. [Section 5] The definition of g(μ̂) is given as (1/p) Σ (X̄_Ti + X̄_Ri)^2, but the context (Eq. [3]) and the f2 definition require (X̄_Ti − X̄_Ri)^2. This appears to be a sign typo and should be corrected.
  3. [Equation [2]] Equation [2] is described in text but not explicitly displayed as a numbered equation; the reader cannot verify the form of h(σ²) or the Taylor approximation. Display the full equation.
  4. [Section 7] Notation is inconsistent: the text alternates between h(σ̂)/n and h(σ̂²)/n. Since h is a function of variances, h(σ̂²)/n is correct.
  5. [Section 10] The survey results are presented as narrative; the response rate, question wording, and analysis method are not given. The paper should either present the survey formally (e.g., as supplementary material) or clearly label it as an informal internal poll so readers can weigh its evidentiary value appropriately.
  6. [Reference list] The entry for the South Korean guideline should list the issuing authority properly (e.g., 'South Korean Ministry of Food and Drug Safety. (2023)') or follow the journal's reference style for agency documents.

Circularity Check

0 steps flagged

No significant circularity: the rejection of expected f2 is benchmarked externally against Shah et al., Liu et al., and the Equation [1] hypotheses; only a minor non-load-bearing self-citation to Hoffelder et al. (2022) appears.

full rationale

The paper's derivation chain is not circular. Its central claim — that f̂2,exp adds h(σ̂²)/n into the logarithm instead of subtracting it, thereby moving the statistic further away from the population f2 — is a direct comparison with Shah et al.'s approximation of E(f̂2) (Equation [2]) and Shah's 'intuitive bias correction' (Equation [3]). That comparison is exact in direction and does not use f̂2,exp to define its own target; the target is the population quantity f2 = 100 − 25 log10(1+QMD²) stated in Section 4 and the equivalence hypotheses in Equation [1]. The power/behavioral claim (f̂2,exp can fall below 50 when QMD=0) is a mathematical property of the published formula and is illustrated with a simulation external to the derivation. The endorsement by FDA statisticians Liu et al. (2024) is an independent external benchmark with overlapping subject matter but not authored by this working group. The only self-touching elements are (a) the approving citation of Hoffelder et al. (2022), co-authored by a contributor, as a sound alternative method, and (b) the internal member survey reporting 'no scientifically meaningful advantage.' Neither is load-bearing: the survey is opinion-gathering, not derivation, and the Hoffelder citation appears only in the concluding remark about alternatives, not in the mathematical argument against f̂2,exp. A possible weakness — that the critique presupposes f̂2,exp is meant to estimate population f2 rather than being a deliberately conservative decision statistic — is an estimand/interpretation disagreement, not a circular reduction; the paper anchors its reading in the equivalence hypotheses and in Liu et al.'s statement that the expected f2 estimators 'correct the bias to the opposite direction.' Overall, no specific circular step can be quoted, so a low score is appropriate.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The paper's central argument rests on regulatory-domain assumptions (the QMD < 10 equivalence boundary parameterized by population f2), on Shah et al.'s approximation as the reference for bias, and on documentary claims about other papers' content. It introduces no new entities and no fitted constants; the only hand-chosen inputs are the undisclosed simulation settings for the illustrative example.

free parameters (1)
  • Undisclosed simulation settings for Figure [1] (n, p, variance magnitudes, bootstrap replications) = not disclosed
    The illustrative claim that power is near zero for identical profiles, and the specific lower bootstrap bounds 35.2/43.3, depend on hand-chosen simulation inputs that are not reported in the text.
axioms (4)
  • domain assumption The regulatory equivalence question is parameterized by the population quantity f2 = 100 - 25·log10(1 + QMD²), i.e., QMD < 10, so estimators must be judged by bias and power against this parameter.
    Section 4 and Equation [1]. The entire 'wrong direction of bias correction' critique presupposes that the estimand of interest is the population f2, not a variance-inflated 'expected observed f2' quantity.
  • standard math Shah et al.'s approximation E(f̂2) ≈ 100 - 25·log10(1 + g(μ) + h(σ²)/n) is an adequate reference for the quantitative bias claim.
    Section 5, Equation [2]. The sign argument does not depend on the approximation because h/n > 0 always lowers the estimate, but the magnitude of the claimed bias increase does.
  • domain assumption Equal REF/TEST sample sizes and the variance decomposition (σRi² + σTi²)/n for the mean difference hold; the method under review is only defined for equal n.
    Section 3 and Section 7. The variance-term analysis assumes this structure, which the method itself assumes, and excludes unequal-n cases in practice.
  • domain assumption The reproduced screenshots of Noce et al. (2020) Eq. (3) accurately represent the published formula and its bracket structure.
    Section 6, Equation [4]. The bracket-ambiguity finding is documentary and cannot be re-verified within the manuscript text alone.

pith-pipeline@v1.3.0-alltime-deepseek · 9443 in / 19138 out tokens · 171069 ms · 2026-08-01T20:02:04.172044+00:00 · methodology

0 comments
read the original abstract

The method called 'expected $f_2$' ($\hat{f}_{2,\exp}$), as proposed by Noce et al. (2020) and Xu et al. (2021), has been adopted in two health authority guidelines for dissolution profile comparison when variability precludes the use of the conventional similarity factor $\hat{f}_2$. This position paper, developed by a working group of the European Federation of Statisticians in the Pharmaceutical Industry CMC Statistical Network Europe Special Interest Group (EFSPI CMCSNE SIG), presents a critical evaluation of this method. Fundamental concerns are identified. First, the formula for $\hat{f}_{2,\exp}$ has no traceable origin in the references cited by its proponents. Noce et al. (2020) and Xu et al. (2021) attribute $\hat{f}_{2,\exp}$ to Shah et al. (1998) and Ma et al. (1999, 2000), but neither mentions nor suggests it. Second, no mathematical justification has been provided for the formula. Where Shah et al. (1998) subtract a variance term to reduce the upward bias of $\hat{f}_2$, the $\hat{f}_{2,\exp}$ formula adds this term, thereby increasing rather than correcting the bias. This has also been noted by FDA statisticians Liu et al. (2024). Third, the method exhibits poor statistical properties: for highly variable profiles, the variance term dominates the statistic, resulting in low power even as the true difference between profiles approaches zero. The method can reject equivalence when profiles are identical. Fourth, the formula as published by Noce et al. (2020) contains a notation ambiguity that renders the intended grouping of terms unclear. This ambiguity has propagated into regulatory guidance. A survey of working group members, designed to elicit arguments both for and against the method, found no scientifically meaningful advantage. The EFSPI CMCSNE SIG concludes that $\hat{f}_{2,\exp}$ should not be recommended for dissolution profile comparison.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

5 extracted references

  1. [1]

    This property is claimed by Noce et al

    The T1E rate is, of course, lower than the T1E rate of other approaches. This property is claimed by Noce et al. (2020) and Xu et al. (2021) as an advantage

  2. [2]

    intuitive bias correction

    is the vector of variances. The auxiliary functions g() and h() are introduced to simplify the formulae and to show the relationships between the suggestions of Shah et al. (1998) and the suggestions of the authors on 𝑓̂2,exp (see next page). Under the approximation shown in Equation [2], the bias depends on h(𝝈𝟐)/n and the point estimate 𝑓2̂ is asymptoti...

  3. [3]

    , and the sample-level quantity 𝑓̂2,exp as computed by Noce et al. (2020). 𝑓̂2,exp is obtained by replacing 𝝁 and 𝝈𝟐 with their sample counterparts in the formula for the approximated expected value 𝔼(𝑓2̂ ) (Equation [4]). Equation [4]: Formula as published in Noce et al. (2020, Eq. (3)), retaining the notation used in that work. The orphaned closing brac...

  4. [5]

    expected f2

    because the notations ‘E(𝑓2)’ or 'expected 𝑓2' as used by Noce et al. (2020) and Xu et al. (2021) are misleading. The latter notations falsely suggest that a computable sample statistic White Paper - EFSPI CMCSNE SIG position on the ‘Expected 𝑓2’ – 2026-06-01, Page 9 is an expected value, when in fact 𝔼(𝑓2̂ ) as presented in Equation [2] is a theoretical ...

  5. [6]

    Statistical consequences of the variance term reversal

    For highly variable profiles, the arbitrary term h(𝝈̂)/n becomes dominant in the decision-making process. As a consequence, the power of the 𝑓̂2,exp method will not increase above a certain limit for highly variable profiles even if the true differences between the profile means tend to zero. This behaviour is not a property of a well- constructed test st...