REVIEW 3 major objections 8 minor 2 references
Equivalence testing in pesticide risk assessment -- Evaluation and practical guidance for design, analysis and interpretation
T0 review · 3 major / 8 minor · reviewed 2026-07-09 · glm-5.2
Pith's one-line read Equivalence Test Protects Against False 'Low-Risk' Pesticide Calls
desk verdict The EFSA equivalence test is more protective than the Hotopp test, and covariate adjustment can cut site requirements by ~10 sites — both claims are well-supported by simulation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Equivalence testing (non-inferiority testing against a specific protection goal threshold), anticlustering randomisation (a design-time technique that maximises within-group heterogeneity in covariates so treatment groups are balanced before the experiment begins), and parametric bootstrap simulation calibrated from a regulatory honeybee field study.
What would settle it
If a regulatory-scale field study with 15+ sites per treatment showed variance components substantially larger than those derived from the Rolke et al. (2016) dataset, the power curves and site-count recommendations in this paper would shift upward, potentially making the EFSA framework less feasible than the authors claim.
Extended reading notes
Core claim
The modified equivalence test proposed as a replacement for EFSA's standard test does not improve discrimination between safe and harmful pesticides; it merely shifts the effective significance level upward, allowing a substantial fraction of pesticides exceeding the 10% harm threshold to be falsely classified as 'low risk.' Meanwhile, balancing colonies across treatments using anticlustering randomisation before the experiment, or statistically adjusting for initial colony size, reduces required site counts by 2-3 sites for harmless pesticides and by about 10 sites for pesticides with 5% effects, making the EFSA framework practical at realistic study sizes.
Load-bearing premise
The simulations assume that the variance in honeybee colony measurements observed in a single six-site field study remains constant as studies scale up to 20 or more sites across wider geographic regions. If larger studies encounter more variable landscapes or less standardised colonies, the estimated site requirements would be too low.
Editorial extensions
If this is right
- Regulatory bodies adopting the modified equivalence test would unknowingly permit higher rates of harmful pesticide approval, since the test's nominal significance level no longer controls the false-trust rate at the protection-goal threshold.
- Pre-experimental covariate balancing via anticlustering could become standard practice in ecotoxicology field studies, reducing required replication and costs without compromising statistical rigor.
- The principle of testing against a protection goal from both sides—demonstrating 'low risk' or 'high risk' rather than treating non-significance as either conclusion—could reshape how regulatory agencies interpret underpowered studies.
- The finding that site requirements scale with true effect size aligns regulatory testing effort with expected risk, concentrating expensive large-scale studies on substances most likely to cause marginal harm.
Reading between the lines
- The calibration mismatch the authors identify in the modified test suggests a general principle: any equivalence test that shifts the comparison baseline to a confidence interval bound rather than a point estimate will decouple the nominal significance level from the actual false-trust rate, and this effect will be worse for studies with smaller sample sizes or higher variance.
- If variance components do increase with study scale—as the authors acknowledge is possible when studies span wider regions—the site counts estimated here are lower bounds, and the practical gap between the two tests' error rates could widen further because the modified test's false-trust rate depends on control-group precision.
- The anticlustering approach could be extended to balance multiple pre-treatment covariates simultaneously (e.g., brood cells, pathogen loads, field characteristics), potentially yielding power gains beyond what single-covariate adjustment achieves, though the marginal benefit of each additional covariate would depend on its correlation with the outcome.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript evaluates two equivalence testing frameworks for pesticide risk assessment on honeybees: the EFSA-recommended non-inferiority test and the Hotopp et al. (2024) modified test that shifts the baseline to the lower confidence bound of the control group. Using parametric bootstrap simulations (5,000 iterations) based on variance components from a single regulatory field study (Rolke et al., 2016), the authors show that the two tests share the same false-trust/false-mistrust trade-off (Pareto frontier) but are calibrated differently: the Hotopp test at nominal α=0.2 has a much higher effective α than the EFSA test. The authors then estimate site replication requirements under the EFSA equivalence test for effect sizes of 0% and -5%, demonstrating that covariate adjustment (via anticlustering randomisation or model terms) reduces required replication by roughly 2–10 sites. Practical R functions and interpretive guidance are provided.
Significance. The paper addresses a timely regulatory question with direct policy implications. The central comparative claim—that the Hotopp test inflates the effective false-trust rate relative to its nominal α—is demonstrated via a structural argument (§4.1) and confirmed by simulation (Fig. 2A–B). The external calibration of site requirement estimates against EFSA's independent calculations (34 vs. 36 sites for -5% without adjustment; 9 vs. 10 for 0%) is a notable strength. The authors ship reproducible R code on GitHub and provide a new `experimental_allocation()` function for anticlustering randomisation, which is a concrete methodological contribution. The guidance section (§5, Fig. 4, Table 2) is practically useful for regulatory toxicologists.
major comments (3)
- §2.8, Hotopp test definition: The description of how the Hotopp test's confidence interval is constructed is ambiguous. The text states that the adjusted log ratio is calculated as the difference between the pesticide group's estimated log marginal mean and the lower limit of the 90% CI of the control group's estimated log marginal mean, and that 'the lower bound of the 60% confidence interval for this adjusted log ratio, calculated using the standard error of the regular log ratio, exceeded log(0.9).' It is unclear whether the standard error of the adjusted log ratio properly accounts for the variability introduced by using a random quantity (the lower CI bound of the control mean) as the baseline. If the standard error of the 'regular log ratio' is used without adjustment, the Hotopp test's false-trust rates as simulated may not match what would obtain under a correct variance formula.
- §2.9 and Fig. 2A: The Pareto comparison aligns the two tests by varying α from 1% to 50%. The conclusion that 'no test was Pareto-superior' is based on simulations at two effect sizes (0% and -11%) with 10 sites per treatment. The -11% effect size is only 1 percentage point above the -10% SPG threshold. It would strengthen the claim to verify that Pareto equivalence holds at more extreme effect sizes (e.g., -15%, -20%), since the practical concern is precisely about pesticides with effects well above the SPG being misclassified.
- §3.2, Fig. 3, and §4.3: The site requirement estimates for the -5% effect size (24–34 sites depending on adjustment) are derived under the assumption that variance components from a single six-site study (Rolke et al., 2016) remain constant as the number of sites scales to 30+. The authors acknowledge this in §4.3, but the claim that 'site requirements remain lower than those implied by the former EFSA guidance' (abstract, §3.2, §4.2) is load-bearing on these numbers. If variance scales upward with study region size, the relative advantage over the 2013 guidance could narrow or reverse. The external calibration (34 vs. 36 sites) partially mitigates this concern, but only at one effect size and one adjustment strategy. A brief sensitivity analysis or at least a quantitative bound on how large the variance inflation would need to be to change the conclusions would substantially strengthen.
minor comments (8)
- §2.2: The step numbering in the simulation procedure lists 'vi) evaluate the model' followed by 'iv) evaluate the model' — the second 'iv' should be 'vii' or similar.
- Fig. 2 caption: 'Simulations for panel B were done with 10 sites per treatment' appears to contradict the panel B description, which varies the number of sites from 4 to 20. Please clarify.
- Table 1: The 'Required statistical power' row for EFSA GD (2023) states 'Implied.' It would help to state explicitly what is implied (i.e., the α=0.2 equivalence test implicitly requires sufficient power to reject the null at the SPG).
- §4.1: The statement that 'reduced precision in estimating only the control mean increases the probability that a pesticide is classified as low risk' (consequence b) could benefit from a brief explanation of the mechanism, as this is counterintuitive.
- References: Wintermantel, 2026 (experimental_allocation R function, Zenodo) has a future date. Please verify this is correct or update.
- §5: The recommendation to replace 'pre-notification requirements' with 'stricter pre-registration' is a policy suggestion that goes beyond the statistical findings. Consider softening or flagging as an opinion.
- Fig. A5 caption references panels A–D but the figure description does not clearly map which panels correspond to which sample size/timepoint combinations. Please verify panel labels match the caption text.
- §2.8: The sentence beginning 'For the Hotopp equivalence test, the log ratio...' is a long run-on sentence. Consider breaking it up for readability.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive review. All three major comments are addressed below. In brief: (1) we clarify that our implementation of the Hotopp test deliberately uses the standard error of the regular log ratio, matching the procedure as described by Hotopp et al. (2024), and we will add explicit discussion of the variance-propagation issue the referee raises; (2) we will add simulations at -15% and -20% effect sizes to confirm Pareto equivalence at more extreme deviations from the SPG; and (3) we will add a quantitative sensitivity analysis showing how large variance inflation would need to be to overturn our site-requirement conclusions.
read point-by-point responses
-
Referee: §2.8, Hotopp test definition: The description of how the Hotopp test's confidence interval is constructed is ambiguous. The text states that the adjusted log ratio is calculated as the difference between the pesticide group's estimated log marginal mean and the lower limit of the 90% CI of the control group's estimated log marginal mean, and that 'the lower bound of the 60% confidence interval for this adjusted log ratio, calculated using the standard error of the regular log ratio, exceeded log(0.9).' It is unclear whether the standard error of the adjusted log ratio properly accounts for the variability introduced by using a random quantity (the lower CI bound of the control mean) as the baseline. If the standard error of the 'regular log ratio' is used without adjustment, the Hotopp test's false-trust rates as simulated may not match what would obtain under a correct variance formula.
Authors: The referee raises an important point. Our implementation deliberately uses the standard error of the regular log ratio (i.e., the difference between the pesticide group mean and the control group mean) without additional adjustment for the variability of the lower CI bound of the control mean. This is consistent with how the Hotopp test is described in Hotopp et al. (2024): the baseline is shifted to the lower 90% CI bound of the control, but the confidence interval width for the resulting 'adjusted log ratio' is based on the standard error of the treatment–control contrast. We agree that this does not properly propagate the additional uncertainty introduced by using a random quantity (the lower CI bound) as the baseline. This is, in fact, part of our critique: the Hotopp test as proposed does not account for this extra variability, which contributes to its inflated effective false-trust rate. However, we acknowledge that the manuscript does not make this sufficiently explicit. We will revise §2.8 to state clearly that (a) the standard error used is that of the regular log ratio, (b) this matches the Hotopp et al. (2024) procedure, and (c) this means the additional uncertainty from using a random baseline is not propagated, which is one reason the Hotopp test's effective α exceeds its nominal α. We will also note that a test that did properly propagate this uncertainty would have wider confidence intervals and thus a lower false-trust rate—though it would still shift the threshold and thus differ from the EFSA test in ways we describe in §4.1. revision: partial
-
Referee: §2.9 and Fig. 2A: The Pareto comparison aligns the two tests by varying α from 1% to 50%. The conclusion that 'no test was Pareto-superior' is based on simulations at two effect sizes (0% and -11%) with 10 sites per treatment. The -11% effect size is only 1 percentage point above the -10% SPG threshold. It would strengthen the claim to verify that Pareto equivalence holds at more extreme effect sizes (e.g., -15%, -20%), since the practical concern is precisely about pesticides with effects well above the SPG being misclassified.
Authors: We agree that verifying Pareto equivalence at more extreme effect sizes would strengthen the claim. The referee is correct that the -11% effect size is close to the SPG threshold. We will add simulations at -15% and -20% true effect sizes (with 10 sites per treatment, 6 colonies per site) to the Pareto comparison and report the results in a revised Figure 2A (or a supplementary figure). We expect Pareto equivalence to hold because the structural argument in §4.1—the Hotopp test shifts the threshold by approximately half the span of the 90% CI of the control mean, which affects the false-trust rate uniformly across effect sizes—does not depend on the specific effect size being close to the SPG. However, we will verify this empirically and report the results honestly. If any deviation from Pareto equivalence emerges at extreme effect sizes, we will report it and discuss the implications. revision: yes
-
Referee: §3.2, Fig. 3, and §4.3: The site requirement estimates for the -5% effect size (24–34 sites depending on adjustment) are derived under the assumption that variance components from a single six-site study (Rolke et al., 2016) remain constant as the number of sites scales to 30+. The authors acknowledge this in §4.3, but the claim that 'site requirements remain lower than those implied by the former EFSA guidance' (abstract, §3.2, §4.2) is load-bearing on these numbers. If variance scales upward with study region size, the relative advantage over the 2013 guidance could narrow or reverse. The external calibration (34 vs. 36 sites) partially mitigates this concern, but only at one effect size and one adjustment strategy. A brief sensitivity analysis or at least a quantitative bound on how large the variance inflation would need to be to change the conclusions would substantially strengthen.
Authors: This is a fair and important concern. We will add a quantitative sensitivity analysis to §4.3 (or a supplementary section) that computes how much variance inflation would be required to change the key conclusions. Specifically, we will re-run the site-requirement simulations for the -5% effect size under inflated variance components (e.g., 1.25×, 1.5×, 2× the site-level and colony-level variance from Rolke et al.) and report the resulting site requirements under both the EFSA 2023 equivalence test and the 2013 difference test. This will allow readers to see the break-even point at which the advantage of the 2023 guidance over the 2013 guidance narrows or reverses. We note that the comparison between the 2023 and 2013 guidance is somewhat robust to variance inflation because both frameworks are affected by higher variance—the 2013 guidance requires detecting a -7% effect with 80% power via difference testing, while the 2023 guidance requires demonstrating equivalence for a -5% effect. If variance increases, both require more sites, but the relative ordering may or may not be preserved depending on the magnitude. The sensitivity analysis will make this explicit. We will also temper the language in the abstract and §4.2 to make clear that this comparison holds under the variance assumptions stated and subject to the sensitivity analysis. revision: yes
Circularity Check
No significant circularity. The paper's central comparison of equivalence tests rests on structural properties of the test constructions, not on circular logic or self-citation chains.
full rationale
The paper's central claims are derived from independent simulation of two structurally different statistical tests (EFSA equivalence test vs. Hotopp modified equivalence test). The key result — that the EFSA test is more protective against false 'low-risk' classifications — follows from the mechanical property that the Hotopp test shifts the comparison baseline to the lower 90% CI bound of the control mean, which inflates the effective α relative to the nominal level. This is demonstrated via Pareto analysis (Fig. 2A) and false-trust curves (Fig. 2B) and does not reduce to a fitted input by construction. The variance components used in simulations are derived from a single external dataset (Rolke et al., 2016), and the resulting site requirement estimates are externally validated against EFSA's own independent estimates (34 vs. 36 sites for -5% effect; 9 vs. 10 for 0% effect), providing calibration rather than circularity. The only self-citation is Wintermantel (2026) for the R function experimental_allocation(), which is a practical tool and does not bear on the central comparative claims. The DHARMa package citation (Hartig, 2024) is for model validation, not a load-bearing premise. No step in the derivation chain reduces to its own inputs by definition or construction.
Assumptions & free parameters
free parameters (3)
- Effect sizes simulated =
0%, -5%, -10% to -20%
- Alpha levels =
0.2 (main), 0.01-0.5 (comparison)
- Number of sites and colonies =
4-45 sites, 6 colonies
assumptions (3)
- domain assumption Variance components derived from the Rolke et al. (2016) control data are representative of typical honeybee field studies and remain constant as study size increases.
- domain assumption A negative binomial GLMM with log-transformed initial bee count as a fixed effect and colony nested in site as random effects adequately captures the data-generating process.
- standard math The EFSA specific protection goal (SPG) of a 10% reduction in colony size is the correct regulatory threshold.
Cite this review
Pith. "Pith review of Equivalence testing in pesticide risk assessment -- Evaluation and practical guidance for design, analysis and interpretation." pith.science (2026). https://pith.science/paper/5EEKFFUR
@misc{pith2026260707543,
author = {Pith},
title = {Pith review of: Equivalence testing in pesticide risk assessment -- Evaluation and practical guidance for design, analysis and interpretation},
year = {2026},
howpublished = {\url{https://pith.science/paper/5EEKFFUR}},
note = {Machine review of arXiv:2607.07543}
}
read the original abstract
Harmful pesticide effects exceeding specific protection goals (SPG) may go undetected in underpowered experimental designs. Regulatory honeybee field studies have consistently failed to reach the statistical power required under European Food Safety Authority (EFSA) guidance, which may have caused approval of high-risk substances. Therefore, EFSA advised a shift from testing the null hypothesis of 'no effect' to equivalence testing. Under this approach, a pesticide is classified as 'low risk' if the null hypothesis that its effect exceeds the SPG can be rejected. For honeybees, the recommended SPG is a colony size reduction below 10%. Critics have argued that this framework requires excessive site replication to demonstrate pesticide safety and proposed an alternative equivalence test defining treatment effects relative to the lower bound of the 90%-control-group confidence interval. Using simulations mimicking a regulatory honeybee field study, we show that although the two equivalence tests share the same trade-off between false 'low-risk' and false 'high-risk' classifications, only EFSA's original recommendation reliably identifies pesticides with effects > SPG at alpha = 0.2. Our results show that increasing site replication beyond the current practice is unavoidable for a reliable regulatory assessment. However, for pesticides with effect sizes of 5% or less, site requirements remain lower than those implied by the power requirement of the former EFSA guidance. Moreover, covariate adjustment through a model term or balanced colony allocation using anticlustering randomisation can reduce site requirements without losing power and thus save costs. Finally, we provide guidance and R functions for anticlustering randomisation and equivalence testing for pesticide risk assessment.
Figures
Reference graph
Works this paper leans on
-
[1]
doParallel: Foreach Parallel Adaptor for the “parallel” Package. https://doi.org/10.32614/CRAN.package.doParallel Microsoft, Weston, S.,
-
[2]
https://doi.org/10.32614/CRAN.package.foreach 2 Fig
foreach: Provides Foreach Looping Construct. https://doi.org/10.32614/CRAN.package.foreach 2 Fig. A1. Probability that a pesticide is classified as 'low risk' across different tests and true effect sizes, comparing original and simulated datasets. Each of the 20 curves per panel is calculated based on 500 simulations. Vertical dashed lines indicate the -7...
Reviewed July 9, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.