{"id":"aaa8657d-bfeb-4f23-ad62-2b409d7b89be","arxiv_id":"1908.04695","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Blinded sample size re-estimation raises the type I error rate in equivalence and non-inferiority tests, with peaks around 6.3 percent for small interim samples, and can be controlled by minimum and maximum sample size caps.","lead":"This paper shows that clinical equivalence trials that adjust their patient numbers after a blinded halfway check can end up with false-positive rates above the intended 5 percent. It explains why this happens and gives simple limits on interim sample sizes that keep the error rate close to nominal.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unproven z/t sample-size formula equivalence undermines the quantitative Table 2 and the nMin/nMax recommendations; the qualitative inflation finding itself stands.","rationale":"I read the paper as making two distinct claims: a qualitative one that blinded SSR in equivalence testing can produce non-negligible type I error inflation, and a quantitative one that the inflation pattern and the suggested nMin/nMax caps are as described in Table 2 and Section 6. The qualitative claim is well supported by the variance decomposition in Eq. (2), the conditional rejection analysis in Figure 4, and the confirmatory simulations; the mechanism is coherent and consistent with prior literature. The quantitative claim rests on the unsupported assertion in Section 4.3 that the z-formula and the noncentral-t formula produce identical inflation patterns. The exact t-formula differs from the z-formula most in small samples, precisely the settings where the paper finds the largest inflation, so this is a genuine load-bearing assumption for the recommendations even though it does not overturn the existence of inflation. I also noted minor presentation issues that do not affect the main argument: Eq. (4) is written for σ=1 while retaining a σ^2 factor, and the displayed unconditional expression in Appendix A appears to use the conditional value 0.0015 where the surrounding bracket values indicate the joint value 0.0009 should be used. No code or data were provided, which is another reason the quantitative calibration remains conditional. The reader's weakest assumption identifies the same concern, and the appropriate verdict remains CONDITIONAL; no further movement is warranted.","tokens_in":12912,"tokens_out":13817,"duration_ms":124798,"concrete_test":"Re-run the simulation grid (same ñ ∈ {10,...,80}, nMin/ñ, nMax/ñ, and δ0/σ grid, at least 10^6 replicates per cell) but compute N̂ at each interim by the exact noncentral-t TOST sample-size formula, using an iterative solver such as R's PowerTOST::sampleN.TOST or equivalent, and apply the same nMin/nMax rule. Compare Table 2 peak Case 1 rates and their δ0/σ locations, plus the heatmaps, with the z-formula results. If the peak α values and locations agree within Monte Carlo error (roughly 0.05 percentage points for 10^6 replicates), the concern is resolved; if they differ materially, the quantitative recommendations must be re-derived from the t-based formula.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main gap is the unproven equivalence of the z-based sample-size formula in Eq. (3) and the exact noncentral-t TOST formula, asserted without support in Section 4.3. The exact per-group sample size n solves n = 2(t_{1-β/2,2n-2}+t_{1-α,2n-2})^2 / (δ0/σ̂_T)^2, so the degrees of freedom appear on both sides of the equation; Eq. (3) removes this dependence by using fixed z quantiles. The two mappings from σ̂_T^2 to N̂ therefore differ, and the difference is largest for small ñ and small σ̂_T^2, which is exactly the region where the paper's mechanism operates (Section 2 and Figure 4). Because the SSR affects the final test only through the random second-stage size m = f(σ̂_T^2), replacing f can change both the magnitude of the inflation and the location of its peak along δ0/σ. Table 2, Figure 3, and the Section 6 recommendation that max α can be limited to within 5.3% are all computed only with Eq. (3), so these quantitative claims are not yet supported for the t-distribution-based sample-size calculation used in practice. The qualitative conclusion that blinded SSR inflates type I error is not threatened by this concern; the transfer of the quantitative numbers to practice is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript investigates the type I error behavior of blinded sample size re-estimation (SSR) in two-group equivalence trials analyzed by TOST. Under the null hypothesis with a non-zero margin delta0, the interim total variance estimate sigma_hat_T^2 has expectation sigma^2 times a factor that increases with delta0^2/sigma^2, so a small sigma_hat_T^2 tends to occur for stage-1 data that already favor equivalence. An SSR rule that increases the second-stage sample size with sigma_hat_T^2 therefore preserves favorable data and dilutes unfavorable data, inflating the type I error. The paper derives this mechanism analytically, quantifies it by one-million-run simulations over a grid of stage-1 sample sizes, minimum and maximum sample size caps, and effect sizes, and gives practical recommendations: choose stage-1 sample size at least 15, impose nMin >= 2*nTilde for nTilde <= 30, impose nMax <= 3*nTilde, and claim maximum alpha can be limited to within 5.3%. An appendix provides an analytic numerical evaluation for a threshold-based SSR rule.","tokens_in":13192,"tokens_out":10059,"duration_ms":95899,"significance":"If the quantitative conclusions are supported, this is a useful applied paper: it gives a transparent explanation for a known but under-explained phenomenon, confirms non-negligible inflation for small but realistic stage-1 sizes, and offers concrete design constraints. The paper's strengths include a correct distribution-theoretic derivation of E(sigma_hat_T^2), a coherent four-case decomposition of TOST outcomes, very large simulations with tight Monte Carlo error, and an appendix with exact numerical evaluation of type I error for a threshold rule. The qualitative finding of inflation is robust. The main weakness is that the quantitative peak values and recommendations are computed with a normal-approximation sample size formula whose equivalence to the t-based formula used in practice is only asserted.","major_comments":[{"comment":"The assertion that it is 'essentially irrelevant' whether the z-based formula in Eq. (3) or the noncentral-t-based formula is used is not supported. The t-based sample size formula has degrees of freedom on both sides of the equation, so the mapping from sigma_hat_T^2 to the second-stage size N-hat differs from Eq. (3), and the difference is largest for small nTilde and small sigma_hat_T^2, precisely the regime where the paper's mechanism operates (Section 2 and Figure 4). Because the SSR enters the final test only through N-hat = f(sigma_hat_T^2), a different f can change both the magnitude of the inflation and the location of its peak along delta0/sigma; the claim of 'identical patterns' needs a proof or, more practically, a sensitivity analysis over the same grid using the t-based formula, for example via PowerTOST. Until such a comparison is provided, the numerical values in Table 2 and the practical bound in Section 6 are not established for the t-based calculations used in practice.","section":"Section 4.3, Eq. (3), Table 2"},{"comment":"The headline recommendation that 'maximum alpha can be limited to within 5.3%' is stated without a supporting summary. Table 2 reports peaks only for the uncapped case 0 <= m < infinity, and Figure 3 is a collection of heatmaps from which the reader cannot verify the maximum over the entire grid of settings satisfying the proposed rules (nTilde >= 15, nMin >= 2*nTilde for nTilde <= 30, nMax <= 3*nTilde). Please report the maximum observed Case-1 probability and its Monte Carlo standard error across all grid points in the recommended design class, together with the corresponding (nTilde, nMin, nMax, delta0/sigma) configuration. This is needed to make the 5.3% claim reproducible and to assess its sensitivity to the z/t issue raised above.","section":"Section 6"}],"minor_comments":[{"comment":"The displayed variance formula is dimensionally inconsistent: with sigma not set to 1 it should read Var(sigma_hat_T^2) = 2*sigma^4/(2*nTilde - 1)^2 * [2*nTilde - 1 + nTilde*delta^2/sigma^2], not 2*sigma^2 times the bracket. The simulations set sigma = 1, so this does not affect the numerical results, but the formula should be corrected.","section":"Section 5.1, Eq. (4)"},{"comment":"The sentence 'In each simulated sample t-test is conducted between the two groups' should read 'the TOST procedure, i.e., two one-sided t-tests, is conducted', since the type I error rates in Table 2 and Figure 3 are for the equivalence test.","section":"Section 4.2"},{"comment":"The phrase 'If the true delta > 0' should be phrased as 'if the null configuration delta = delta0 > 0 holds', because delta is the fixed parameter and the sentence is otherwise easy to misread.","section":"Section 2"},{"comment":"The text describes bars as 'percentage of rejections which occur in intervals' of sigma_hat_T^2; it would help to state explicitly that these are joint proportions (rejection and sigma_hat_T^2 interval), not conditional rejection rates, because the subsequent comparison of gray and black bars depends on this.","section":"Section 5.1, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an applied statistics journal and the qualitative finding is likely correct. I would not reject over the z/t issue because the qualitative result is supported by the derivation and simulations, and a sensitivity analysis is feasible. The authors should be asked to either prove or empirically test the z/t equivalence and to supply a summary table supporting the 5.3% bound. The self-citation [4] is not load-bearing for the paper's claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the core finding holds up: blinded SSR in equivalence and non-inferiority testing does inflate the type I error rate, and the paper gives the mechanism cleanly. Under a shifted null, the expected total variance at the interim depends on the squared true difference, so a small total variance estimate is informative against H0. An SSR rule that recruits more when the variance estimate is large preserves favorable data and dilutes unfavorable data — that is the real engine of the inflation. Second, the paper's quantitative recommendations, including Table 2 and the 5.3% cap, rest on an assertion in Section 4.3 that the z-based sample size formula and the exact noncentral t formula give essentially identical inflation patterns. That assertion is not demonstrated, and the t-based formula has degrees of freedom on both sides of the equation, so the mapping from the interim variance to the final sample size genuinely differs. The difference is largest in the small-sample, small-variance region where the paper's mechanism operates. The qualitative conclusion is not threatened, but the practical numbers may be.\n\nThe derivation in Equation (2) is correct and is a genuinely useful addition to the earlier work by Friede and Kieser and by Lu. The simulations are extensive — one million runs per setting give tight Monte Carlo confidence bands — and the heatmaps in Figure 3 make the dependence on stage-1 sample size, nMin, and nMax easy to see. Appendix A, with an exact numerical evaluation for a threshold rule, is a nice complement to the simulation work.\n\nSoft spots, in proportion. The z/t equivalence is the main one. It is load-bearing for the quantitative claims, and as written it is an unsupported assertion. A sensitivity table using the t-based formula, or even a small side-by-side simulation, would settle it. I would not call this fatal, because the direction and rough magnitude of the inflation are consistent with prior literature and the mechanism is sound, but the practical recommendations overreach as written. Equation (4) also has a dimensional slip; with sigma set to 1 it does not affect the numbers, but it should be fixed. No code or data are provided, which is a missed opportunity — the simulation design is simple enough that shipping a script would let any reader confirm the z/t question directly.\n\nThis paper is for applied statisticians in pharma and biostatistics who design equivalence or non-inferiority trials and might be tempted to use blinded SSR. It is worth their time. I would send it to a serious referee rather than desk reject it; the right referee will ask for the z/t sensitivity analysis, and if that is added the paper will be in good shape. I would cite it, mainly for the mechanism.","headline":"Solid mechanism and extensive simulations, but the quantitative guardrails rest on an unverified z/t sample-size equivalence; the qualitative inflation result stands.","tokens_in":13776,"tokens_out":4094,"would_cite":true,"duration_ms":39500,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Blinded sample size re-estimation in equivalence testing can inflate type I error rates above nominal levels, with the mechanism traced to the total variance estimate's dependence on the squared mean difference.","keywords":["bioequivalence","biosimilar","non-inferiority","sample size re-estimation","TOST","type I error control","blinded interim analysis","equivalence testing"],"falsifier":"Rerun the simulation grid using the exact noncentral t-distribution sample size formula instead of Equation (3) and compare the peak inflation values in Table 2; if the exact formula shifts the peaks by more than simulation error for small interim samples, the paper's practical caps on interim size, minimum total, and maximum total would not transfer as stated.","tokens_in":12665,"feed_emoji":"📊","tokens_out":7673,"duration_ms":68696,"temperature":0.7,"pith_summary":"Blinded sample size re-estimation (SSR) is a practical tool for adjusting a trial's size after an interim look without unblinding treatment groups. This paper tries to establish that in equivalence testing, that tool breaks the nominal type I error guarantee: the false-equivalence rate can climb well above 5%, reaching about 6.3% in simulations with 10 subjects per group at the interim. The explanation is structural, not a numerical quirk. Under the null hypothesis with a positive equivalence margin, the one-sample total variance estimate has expectation inflated by the squared mean difference, so a small observed variance tends to coincide with data favoring equivalence. An SSR rule that recruits more subjects when the variance looks large therefore preserves favorable data and dilutes unfavorable data, and the paper shows by simulation and an exact appendix calculation that the resulting inflation is non-negligible across practically relevant settings.","feed_headline":"Blinded sample size updates inflate equivalence-test error rates","feed_subtitle":"Small interim samples can push the false-equivalence rate to 6.3%; here is why and how to cap it.","key_machinery":"The load-bearing object is the blinded total variance estimate $\\hat{\\sigma}_T^2$, the one-sample variance of all observations pooled across the two treatment groups. Its expectation under the shifted null (Equation (2)) is the identity that carries the argument: it contains a positive term proportional to $\\delta^2/\\sigma^2$, so the estimator is not unbiased under $H_0$ when the margin $\\delta_0$ is positive. This makes small observed variances coincide with data in favor of the alternative. The sample size rule then turns that correlation into error inflation because the planned second-stage size $\\hat{N}$ is an increasing function of $\\hat{\\sigma}_T^2$; the appendix shows that an exact noncentral chi-square decomposition allows the type I error to be computed for a threshold rule.","core_discovery":"The paper's central claim is that the type I error of TOST equivalence testing is not preserved by blinded SSR, and that the violation is driven by Equation (2): under $H_0$ with $\\delta=\\delta_0>0$, $$E(\\hat{\\$\\sigma$}$_T^{2}$)=\\$sigma^{2}$\\left(1+\\frac{\\tilde{n}_1\\tilde{n}_2}{\\tilde{n}_*(\\tilde{n}_*-1)}\\frac{\\$delta^{2}$}{\\$sigma^{2}$}\\right).$$ Because the total variance estimator is an increasing function of the squared true mean difference, small values of $\\hat{\\sigma}_T^2$ are evidence against $H_0$. A blinded SSR rule whose final sample size is an increasing function of $\\hat{\\sigma}_T^2$ therefore tends to stop early when the data already favor rejection and to enlarge the study when the data do not, inflating the chance of declaring equivalence. Simulations with one million replications per setting show peak type I error rates of 6.26% at $\\tilde{n}=10$ per group and 5.23% at $\\tilde{n}=60$, with the worst inflation at standardized equivalence margins near $\\delta_0/\\sigma \\approx 0.55$ to $1.20$. The appendix gives an exact numerical evaluation for a threshold rule, confirming inflation analytically rather than only by simulation.","pith_inferences":["The same mechanism would predict $\\alpha$ deflation for any blinded SSR rule in which the second-stage sample size is a decreasing function of $\\hat{\\sigma}_T^2$; this is a direct consequence of Equation (2) and could be tested with the paper's simulation grid.","The qualitative argument likely extends to other blinded variance estimators that are monotone increasing in the squared mean difference, not just the simple total variance estimate; verifying this would require new simulations.","A practical implication the paper leaves implicit is that pre-specifying a narrow $n_{\\rm Min}$–$n_{\\rm Max}$ window is the cheapest fix, but it partially defeats the purpose of an interim reassessment; unblinded SSR is the alternative that restores flexibility while controlling error."],"forward_implications":["With an interim sample of at least 15 subjects per group and a minimum total sample size of at least twice that, the maximum type I error in the simulated settings stays within 5.3%.","The lower bound on the final sample size ($n_{\\rm Min}$) is the main lever for controlling inflation at practically relevant equivalence margins; the upper bound matters mainly when the margin is much smaller than the standard deviation.","The inflation persists even at larger interim samples (about 5.2% at $\\tilde{n}=80$), so it is not only a small-sample curiosity.","A design that keeps the second-stage sample size within a narrow range resembles a fixed design and keeps $\\alpha$ closer to its nominal level, at the cost of reducing the flexibility that motivates SSR."],"supporting_citations":[{"why":"Introduces simple blinded SSR procedures that the paper positions as the practical baseline for the methods under study.","marker":"[2]"},{"why":"Reviews internal pilot designs and blinded SSR, establishing the context in which the type I error question is raised.","marker":"[3]"},{"why":"Shows that the two-sample t statistic after blinded SSR has a non-standard distribution, providing evidence that type I error is not controlled.","marker":"[5]"},{"why":"Documents inflation in non-inferiority and equivalence trials, the direct phenomenon this paper explains and quantifies.","marker":"[6]"},{"why":"Discusses bias versus variance in blinded variance estimation, supporting the mechanism linking variance estimation to error rates.","marker":"[8]"},{"why":"Defines the two one-sided tests (TOST) procedure that the paper uses for equivalence testing.","marker":"[9]"},{"why":"Presents unblinded SSR for equivalence, mentioned as an alternative when blinded SSR is problematic.","marker":"[13]"}],"fun_headline_variants":["Equivalence test error rates spike with blinded sample re-estimation","Blinded re-estimation breaks equivalence test type I error control","False equivalence risk rises to 6.3% from blinded interim sizing","Blinded sample size updates inflate equivalence error to 6.3%","Small blinded interim samples ruin equivalence test reliability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative recommendations rest on the assertion that the simpler normal-approximation sample size formula (Equation (3)) produces the same type I error inflation pattern as the exact noncentral t-based formula; the paper states this is essentially irrelevant but does not supply a proof or a direct comparison.","fun_headline_variants_meta":{"raw":{"variants":["Equivalence test error rates spike with blinded sample re-estimation","Blinded re-estimation breaks equivalence test type I error control","False equivalence risk rises to 6.3% from blinded interim sizing","Blinded sample size updates inflate equivalence error to 6.3%","Small blinded interim samples ruin equivalence test reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000642,"raw_usage":{"total_tokens":2930,"prompt_tokens":896,"completion_tokens":2034,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1947}},"tokens_in":512,"tokens_out":2034,"duration_ms":14815,"temperature":1.0,"reasoning_tokens":1947,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:35:54.620845+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the simulation grid using the exact noncentral t-distribution sample size formula instead of Equation (3) and compare the peak inflation values in Table 2; if the exact formula shifts the peaks by more than simulation error for small interim samples, the paper's practical caps on interim size, minimum total, and maximum total would not transfer as stated.","supporting_citations":[{"cited_title":"Simple procedures for blinded sample size adjust- ment that do not aﬀect the type I error rate","cited_arxiv_id":null,"evidence_quote":"Introduces simple blinded SSR procedures that the paper positions as the practical baseline for the methods under study."},{"cited_title":"Sample size recalculation in internal pilot study designs: a review","cited_arxiv_id":null,"evidence_quote":"Reviews internal pilot designs and blinded SSR, establishing the context in which the type I error question is raised."},{"cited_title":"Distribution of the two-sample t-test statistic following blinded sample size re-estimation","cited_arxiv_id":null,"evidence_quote":"Shows that the two-sample t statistic after blinded SSR has a non-standard distribution, providing evidence that type I error is not controlled."},{"cited_title":"Blinded sample size reassessment in non-inferiority and equivalence trials","cited_arxiv_id":null,"evidence_quote":"Documents inflation in non-inferiority and equivalence trials, the direct phenomenon this paper explains and quantifies."},{"cited_title":"Blinded sample size re-estimation in superiority and noninferiority trials: bias versus variance in variance estimation","cited_arxiv_id":null,"evidence_quote":"Discusses bias versus variance in blinded variance estimation, supporting the mechanism linking variance estimation to error rates."},{"cited_title":"A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailabil- ity","cited_arxiv_id":null,"evidence_quote":"Defines the two one-sided tests (TOST) procedure that the paper uses for equivalence testing."},{"cited_title":"Controlling the type I error rate in two-stage sequential adaptive designs when testing for average bioe- quivalence","cited_arxiv_id":null,"evidence_quote":"Presents unblinded SSR for equivalence, mentioned as an alternative when blinded SSR is problematic."}],"review_version":1}