REVIEW 3 major objections 4 minor 110 references
Should we test the model assumptions before running a model-based test?
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that, under four explicit conditions, a procedure that tests model assumptions first and then chooses between two tests can have strictly higher power than either test applied unconditionally.
desk verdict A careful, honest survey plus a correct but simple lemma; the practical conclusion overreaches because the key independence relaxation is unproved, but it deserves serious refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the combined test $\Phi_C(z)=\Phi_{\mathrm{MC}}(z)$ if $\Phi_{\mathrm{MS}}(z)=0$, and $\Phi_{\mathrm{AU}}(z)$ otherwise. The proof machinery is a two-point mixture setup where the whole dataset is drawn from $P_\theta$ with probability $\lambda$ and from $Q$ with probability $1-\lambda$, making the powers of all three procedures linear in $\lambda$. Assumptions (I)–(III) say the MC test beats the AU test under $P_\theta$, the AU test beats the MC test under $Q$, and the misspecification test rejects the model more often under $Q$; assumption (IV) is independence of both main rejection events from the misspecification decision under both distributions. Lemma 1 combines these to show that at the crossing point $\lambda^*=\Delta_Q/(\Delta_\theta+\Delta_Q)$, the combined procedure's power exceeds both the MC and AU power by the positive term $\frac{\Delta_\theta\Delta_Q}{\Delta_\theta+\Delta_Q}(\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}})$.
What would settle it
Simulate a two-sample setup where the misspecification test and the main test are strongly dependent, for example the Shapiro-Wilk normality test and Welch's t-test on the same small normal sample, and compute the four conditional-probability deviations from independence in the paper's $\delta$ relaxation. Then check whether at $\lambda^*=\Delta_Q/(\Delta_\theta+\Delta_Q)$ the combined procedure's power still exceeds both unconditional powers; if it does not, condition (IV) is essential and the promised relaxation fails in that case.
Extended reading notes
Core claim
The central claim is that preliminary misspecification testing—checking model assumptions before a model-based test—can be beneficial, contrary to much of the existing literature. The authors formalize a two-regime setup: data come either from a distribution $P_\theta$ satisfying the model assumptions (where the constrained MC test has higher power) or from $Q$ violating them (where the unconstrained AU test has higher power), with $\lambda$ the probability of the first regime. Under assumptions (I)–(IV), Lemma 1 establishes that there is a $\lambda^*\in(0,1)$ such that the combined procedure has strictly higher power than both unconditional tests at that mixture. The advantage over the better of the two unconditional tests equals $\frac{\Delta_\theta\Delta_Q}{\Delta_\theta+\Delta_Q}(\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}})>0$, where $\Delta_\theta$ and $\Delta_Q$ are the power gaps and $\alpha_{\mathrm{MS}}$, $\alpha^*_{\mathrm{MS}}$ are the misspecification test's rejection probabilities under $P_\theta$ and $Q$. Section 7 broadens this into the claim that model checking is more useful than the literature suggests, provided conditions (a)–(d) hold.
Load-bearing premise
The load-bearing premise is the independence of the misspecification test from both main tests (assumption (IV)); the authors call it very restrictive and unrealistic in most situations, and their promised relaxation to approximate independence with a small $\delta$ is stated without proof.
Editorial extensions
If this is right
- A combined procedure can strictly outperform both of its constituent tests in power for some mixture of valid and violated model assumptions, not just match the better one.
- The conditions specify what makes model checking worthwhile: an informative misspecification test, approximate independence between the misspecification test and the main tests, and clear power advantages of both main tests in their own regimes.
- Model checking should be aimed at detecting violations that are problematic for the main test, not at verifying that assumptions hold exactly.
- Evaluating combined procedures only at $\lambda=0$ or $\lambda=1$, as most of the literature does, is too pessimistic; the mixture perspective can be more favorable.
- In the two-sample normality example, the combined procedure performs close to the better of the t-test and the Wilcoxon-Mann-Whitney test, and can even exceed both in power for a skew-normal distribution, though its type 1 error can be mildly anti-conservative.
Reading between the lines
- Extension: A practical corollary the authors leave implicit is that the misspecification test's significance level could be tuned to maximize the positive advantage term, since $\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}}$ enters linearly; higher misspecification-test levels may be better when the cost of failing to switch to the AU test is large.
- Extension: The $\lambda$-mixture idea can be carried further: with a prior over a continuous family of violations, the same linearity argument would give conditions for the combined procedure to beat 'always use AU' or 'always use MC' in average power.
- Extension: Because assumption (IV) is known to fail in common designs such as crossover trials, a quantitative bound on the power loss as the dependence parameter $\delta$ grows is needed; the paper promises but does not deliver this, so the practical message currently rests on a plausibility argument.
- Extension: The framework suggests a diagnostic for practice: before adopting a check-then-test protocol, simulate the two regimes and the misspecification test's sensitivity to see whether the four conditions hold for the specific tests and plausible violation distributions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether model assumptions should be tested before running a model-based test, and argues that the literature is too pessimistic about such 'combined procedures.' It formalizes a protocol involving a misspecification (MS) test, a model-based constrained (MC) test, and an alternative unconstrained (AU) test, reviews the existing literature, and presents a new theoretical result. Lemma 1 states that under four assumptions—including independence of the MS rejection event from the MC and AU rejection events—there exists a mixture weight λ in (0,1) such that the combined procedure has strictly higher power than either the MC or the AU test. The paper also provides a simulation example (Table 1) and a deferred simulation summarized by Figure 1, and concludes with conditions (a)-(d) under which model checking is worthwhile.
Significance. The paper's main value is conceptual and synthetic: it frames the debate over preliminary model testing in a general way, identifies the needed ingredients for a successful combined procedure, and gives a simple algebraic result showing that under explicit assumptions a combined procedure can beat both unconditional tests. The lemma itself is correct under its stated assumptions and is a useful counterpoint to the uniformly negative conclusions in parts of the literature. However, the practical message depends on an unproved relaxation of the restrictive independence assumption (IV), and the empirical support is partial and partly deferred. If the relaxation can be proved or convincingly demonstrated by simulation, the paper would make a solid contribution; as it stands, the central claim is defensible but not fully established.
major comments (3)
- [Section 6, Lemma 1 and the paragraph following it] The practical conclusion of the paper rests on relaxing assumption (IV), the independence of RMC and RAU from RMS, which the authors themselves call 'very restrictive' and 'unrealistic in most situations.' The relaxation to approximate independence is asserted only in the sentence 'it can be relaxed ... at the price of a more tedious proof that we do not present here,' with no statement of the required bound on δ in terms of Δθ, ΔQ, αMS, and α*MS. Since the proof of Lemma 1 uses exact independence to replace conditional rejection probabilities by unconditional ones, any violation of (IV) enters directly into the sign of Pλ(RC) - Pλ(RMC). Without a quantified condition or a numerical demonstration that the power advantage survives realistic dependence, the central claim that model checking is more useful than the literature suggests is not established.
- [Section 6, Figure 1 and Example 1/Table 1] The claim that the range of λ for which the combined procedure is best is 'quite large' is supported only by simulations that are 'published elsewhere' and by Figure 1, which is described as a typical pattern without reporting sample sizes, numbers of replications, or other methodological details. The only numerical evidence in the paper, Table 1, shows the combined procedure to be anti-conservative for the normal, t3, and skew-normal cases (type 1 error probabilities 0.0512, 0.0515, and 0.0531 against the nominal 0.05) and clearly worse than both the t-test and the permutation test for the exponential distribution in terms of type 2 error (0.4849 versus 0.3389 for the t-test). Thus the empirical support for the practical message is partial, and the power advantage in Lemma 1 is not accompanied by a demonstration that the combined procedure respects the nominal level.
- [Section 6, Lemma 1 and Section 7] The positive result is an existence statement: it guarantees some λ in (0,1) with a strict power advantage, but it gives no expression for the length of the interval of such λ, and the proof shows only that the power difference at λ* is positive under exact independence. Even if the independence relaxation were proved, the practical recommendation in Section 7 would require the difference to be positive over a substantial range of λ, which is not established by the lemma. The authors acknowledge that simulations indicate a large range, but those simulations are deferred; as a mathematical result, Lemma 1 alone is too weak to support the broad conclusion that model checking is worthwhile.
minor comments (4)
- [Example 1, text preceding Table 1] There is a typo in the sentence 'The WMW test is is clearly superior to the t-test for the t3- and skew normal distribution'; the word 'is' is repeated.
- [Figure 1 caption] The caption would benefit from reporting the sample sizes and the number of simulation replicates; as it stands, the reader cannot evaluate the stability of the reported pattern.
- [Section 6, paragraph on the mixture setup] The distinction between the mixture over whole datasets and an ordinary mixture model over individual observations is clear, but the sentence 'the setup has a certain Bayesian flavor' could be expanded to explain how a frequentist should interpret the randomness in λ; at present the interpretability of Pλ as a sampling distribution is not fully discussed.
- [References] A few references are incomplete, for example Campbell (2019) is cited as an arXiv preprint with a DOI and no final publication details, and the entries for Campbell and Dean and for Rasch et al. lack page ranges in the reference list.
Circularity Check
No circularity: Lemma 1 is a conditional derivation under explicit assumptions; the unproved δ-relaxation of (IV) is a stated gap, not a circular step.
full rationale
The paper's central derivation is Lemma 1 (Section 6). Its proof is self-contained algebraic manipulation from assumptions (I)-(IV): it defines Δθ, ΔQ, αMS, α*MS, locates λ*=ΔQ/(Δθ+ΔQ), and shows Pλ*(RC)=Pλ*(RMC)+ΔθΔQ/(Δθ+ΔQ)(α*MS−αMS)=Pλ*(RAU)+the same positive term. Nowhere is the conclusion assumed; (I)-(III) are ordering and usefulness conditions, and (IV) is an independence condition, not a restatement of the target inequality. The paper explicitly flags (IV) as 'very restrictive' and 'unrealistic in most situations', and the promised relaxation to small δ is asserted without proof ('at the price of a more tedious proof that we do not present here'). This is a genuine correctness/generalisability gap that the practical message leans on, but it is not circularity: the exact-(IV) Lemma 1 is proven, and the gap is an unquantified continuity step, not an input disguised as an output. Self-citations (Hennig 2007, 2010; Gelman and Hennig 2017) supply terminology and framing ('misspecification paradox', constructivist model view) and are not load-bearing in the proof. The simulation example (Table 1) is external evidence with stated limitations (anti-conservativity; exponential case), not a fit-then-predict scheme. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no known result is relabelled as new. Score 0.
Assumptions & free parameters
assumptions (8)
- domain assumption Probability mixture form: a dataset comes entirely from Pθ with probability λ and entirely from Q with probability 1-λ.
- domain assumption (I) Δθ = Pθ(RMC) - Pθ(RAU) > 0: the constrained test has better power when the model holds.
- domain assumption (II) ΔQ = Q(RAU) - Q(RMC) > 0: the unconstrained test has better power when the model is violated.
- domain assumption (III) α*_MS = Q(RMS) > α_MS = Pθ(RMS): the misspecification test can distinguish the two regimes at least somewhat.
- ad hoc to paper (IV) RMC and RAU are independent of RMS under both Pθ and Q.
- domain assumption There is a well-defined extension M* of the main null hypothesis into the larger model M, with M* ∩ MΘ = MΘ0.
- domain assumption Level and power are the appropriate criteria for comparing procedures.
- standard math Standard probability linearity for mixtures, and the algebra in the proof of Lemma 1.
Cite this review
Pith. "Pith review of Should we test the model assumptions before running a model-based test?." pith.science (2026). https://pith.science/paper/FECCQTCX
@misc{pith2026190802218,
author = {Pith},
title = {Pith review of: Should we test the model assumptions before running a model-based test?},
year = {2026},
howpublished = {\url{https://pith.science/paper/FECCQTCX}},
note = {Machine review of arXiv:1908.02218}
}
read the original abstract
Statistical methods are based on model assumptions, and it is statistical folklore that a method's model assumptions should be checked before applying it. This can be formally done by running one or more misspecification tests of model assumptions before running a method that requires these assumptions; here we focus on model-based tests. A combined test procedure can be defined by specifying a protocol in which first model assumptions are tested and then, conditionally on the outcome, a test is run that requires or does not require the tested assumptions. Although such an approach is often taken in practice, much of the literature that investigated this is surprisingly critical of it. Our aim is to explore conditions under which model checking is advisable or not advisable. For this, we review results regarding such "combined procedures" in the literature, we review and discuss controversial views on the role of model checking in statistics, and we present a general setup in which we can show that preliminary model checking is advantageous, which implies conditions for making model checking worthwhile.
Reference graph
Works this paper leans on
-
[1]
Journal of Transportation Technologies 7:133--147
Abdulhafedh A (2017) How to detect and remove temporal autocorrelation in vehicular crash data. Journal of Transportation Technologies 7:133--147
2017
-
[2]
Journal of Statistical Planning and Inference 88:47--57
Albers W, Boon PC, Kallenberg WC (2000a) The asymptotic behavior of tests for normal means based on a variance pre-test. Journal of Statistical Planning and Inference 88:47--57
2000
-
[3]
Annals of Statistics 28:195--214
Albers W, Boon PC, Kallenberg WC (2000b) Size and power of pretest procedures. Annals of Statistics 28:195--214
2000
-
[4]
Journal of the American Statistical Association 65:1590--1596
Arnold BC (1970) Hypothesis testing incorporating a preliminary test of significance. Journal of the American Statistical Association 65:1590--1596
1970
-
[5]
Annals of Mathematical Statistics 27:1115--1122
Bahadur R, Savage L (1956) The nonexistence of certain statistical procedures in nonparametric problems. Annals of Mathematical Statistics 27:1115--1122
1956
-
[6]
Annals of Mathematical Statistics 15:190--204
Bancroft TA (1944) On biases in estimation due to the use of preliminary tests of significance. Annals of Mathematical Statistics 15:190--204
1944
-
[7]
Biometrics 20:427--442
Bancroft TA (1964) Analysis and inference for incompletely specified models involving the use of preliminary test(s) of significance. Biometrics 20:427--442
1964
-
[8]
International Statistical Review 45:117--127
Bancroft TA, Han C (1977) Inference based on conditional specification: A note and a bibliography. International Statistical Review 45:117--127
1977
Show all 110 references
-
[9]
Mathematical Proceedings of the Cambridge Philosophical Society 31:223--231
Bartlett MS (1935) The effect of non-normality on the t distribution. Mathematical Proceedings of the Cambridge Philosophical Society 31:223--231
1935
-
[10]
Annals of Statistics 41:802--837
Berk R, Brown L, Buja A, Zhang K, Zhao L (2013) Berk, r., brown, l., buja, a., zhang, k., zhao, l. Annals of Statistics 41:802--837
2013
-
[11]
International Journal of Approximate Reasoning 66:53--72
Bickel DR (2015) Inference after checking multiple bayesian models for data conflict and applications to mitigating the influence of rejected priors. International Journal of Approximate Reasoning 66:53--72
2015
-
[12]
arXiv:191100115 (stat), accepted by Methods in Ecology and Evolution doi:https://doi.org/10.1111/2041-210X.13559, ://arxiv.org/abs/1911.00115
Campbell H (2019) The consequences of checking for zero-inflation and overdispersion in the analysis of count data. arXiv:191100115 (stat), accepted by Methods in Ecology and Evolution doi:https://doi.org/10.1111/2041-210X.13559, ://arxiv.org/abs/1911.00115
2019 arXiv
-
[13]
Statistics in Medicine 33(6):1042--1056, doi:https://doi.org/10.1002/sim.6021
Campbell H, Dean C (2014) The consequences of proportional hazards based model selection. Statistics in Medicine 33(6):1042--1056, doi:https://doi.org/10.1002/sim.6021
2014 doi
-
[14]
Journal of the Royal Statistical Society, Series B 158:419--466
Chatfield C (1995) Model uncertainty, data mining and statistical inference (with discussion). Journal of the Royal Statistical Society, Series B 158:419--466
1995
-
[15]
Cambridge University Press, Cambridge
Cox DR (2006) Principles of Statistical Inference. Cambridge University Press, Cambridge
2006
-
[16]
Australian Journal of Statistics 22:143--153
Cressie N (1980) Relaxing assumptions in the one-sample t-test. Australian Journal of Statistics 22:143--153
1980
-
[17]
Chapman & Hall/CRC, Boca Raton FL
Davies PL (2014) Data Analysis and Approximate Models. Chapman & Hall/CRC, Boca Raton FL
2014
-
[18]
Annals of Statistics 16:1390--1420
Donoho D (1988) One-sided inference about functionals of a density. Annals of Statistics 16:1390--1420
1988
-
[19]
Wiley, New York
Dowdy S, Wearden S, Chilko D (2004) Statistics for Research. Wiley, New York
2004
-
[20]
Journal of the Royal Statistical Society, Series B 57:45--97
Draper D (1995) Assessment and propagation of model uncertainty (with discussion). Journal of the Royal Statistical Society, Series B 57:45--97
1995
-
[21]
Technometrics 18:1--9
Easterling RG (1976) Goodness of fit and parameter estimation. Technometrics 18:1--9
1976
-
[22]
Journal of Statistical Computation and Simulation 8:1--11
Easterling RG, Anderson HE (1978) The effect of preliminary normality goodness of fit tests on subsequent inference. Journal of Statistical Computation and Simulation 8:1--11
1978
-
[23]
Journal of the American Statistical Association 109:991--1007
Efron B (2014) Estimation and accuracy after model selection. Journal of the American Statistical Association 109:991--1007
2014
-
[24]
Journal of Statistical Computation and Simulation 76:803--816
Farrell PJ, Rogers-Stewart K (2006) Comprehensive study of tests for normality and symmetry: extending the spiegelhalter test. Journal of Statistical Computation and Simulation 76:803--816
2006
-
[25]
Statistics Surveys 4:1--39
Fay MP, Proschan MA (2010) Wilcoxon-mann-whitney or t-test? on assumptions for hypothesis tests and multiple interpretations of decision rules. Statistics Surveys 4:1--39
2010
-
[26]
Wiley, New York
de Finetti B (1974) Theory of Probability. Wiley, New York
1974
-
[27]
Econometrica: Journal of the Econometric Society 29:139--170
Fisher FM (1961) On the cost of approximate specification in simultaneous equation estimation. Econometrica: Journal of the Econometric Society 29:139--170
1961
-
[28]
Philosophical Transactions of the Royal Society of London A 22:309--368
Fisher RA (1922) On the mathematical foundation of theoretical statistics. Philosophical Transactions of the Royal Society of London A 22:309--368
1922
-
[29]
Statistics in Medicine 8:1421--1432
Freeman P (1989) The performance of the two-stage analysis of two-treatment, two-period cross-over trials. Statistics in Medicine 8:1421--1432
1989
-
[30]
BMC Dermatology 2:163--174
Gambichler T, Bader A, Vojvodic M, Bechara FG, Sauermann K, Altmeyer P, Hoffmann K (2002) Impact of uva exposure on psychological parameters and circulating serotonin and melatonin. BMC Dermatology 2:163--174
2002
-
[31]
Gamble C, Krishan A, Stocken D, Lewis S, Juszczak E, Doré C, Williamson PR, Altman DG, Montgomery A, Lim P, Berlin J, Senn S, Day S, Barbachano Y, Loder E (2017) Guidelines for the Content of Statistical Analysis Plans in Clinical Trials . JAMA 318(23):2337--2343, doi:10.1001/...
2017
-
[32]
Communications in Statistics - Simulation and Computation 10:163--174
Gans DJ (1981) Use of a preliminary test in comparing two sample means. Communications in Statistics - Simulation and Computation 10:163--174
1981
-
[33]
Journal of the Royal Statistical Society, Series A 180:967--1033
Gelman A, Hennig C (2017) Beyond subjective and objective in statistics (with discussion). Journal of the Royal Statistical Society, Series A 180:967--1033
2017
-
[34]
The American Statistician 102:460--465
Gelman A, Loken E (2014) The statistical crisis in science. The American Statistician 102:460--465
2014
-
[35]
British Journal of Mathematical and Statistical Psychology 66:8--38
Gelman A, Shalizi CR (2013) Philosophy and the practice of bayesian statistics. British Journal of Mathematical and Statistical Psychology 66:8--38
2013
-
[36]
Journal of Economic Surveys 7:145--197
Giles DEA, Giles JA (1993) Pre-test estimation and testing in econometrics: Recent developments. Journal of Economic Surveys 7:145--197
1993
-
[37]
The Lagrange Multiplier principle and other applications
Godfrey LG (1988) Misspecification tests in econometrics. The Lagrange Multiplier principle and other applications. Cambridge University Press, Cambridge
1988
-
[38]
Journal of Statistical Planning and Inference 49:241--260
Godfrey LG (1996) Misspecification tests and their uses in econometrics. Journal of Statistical Planning and Inference 49:241--260
1996
-
[39]
Biometrics 21:469--480 (Corrigendum in Biometrics, 30, 727, 1974)
Grizzle JE (1967) The two-period change-over design and its use in clinical trials. Biometrics 21:469--480 (Corrigendum in Biometrics, 30, 727, 1974)
1967
-
[40]
Journal of the Indian Statistical Association 7:26--29
Gupta VP, Srivastava VK (1993) Upper bound for the size of a test procedure using preliminary tests of significance. Journal of the Indian Statistical Association 7:26--29
1993
-
[41]
Biometrika 49:403--417
Gurland J, McCullough R (1962) Testing equality of means after a preliminary test of equality of variances. Biometrika 49:403--417
1962
-
[42]
Wiley, New York
Hampel FR, Ronchetti EM, Rousseeuw PJ, Stahel WA (1986) Robust Statistics. Wiley, New York
1986
-
[43]
Frontiers in Neuroscience 14:687
Hasler G, Suker S, Schoretsanitis G, Mihov Y (2020) Sustained improvement of negative self-schema after a single ketamine infusion: An open-label study. Frontiers in Neuroscience 14:687
2020
-
[44]
Journal of the American Statistical Association 85:446--452
He X, Simpson DG, Portnoy SL (1990) Breakdown robustness of tests. Journal of the American Statistical Association 85:446--452
1990
-
[45]
MIT Press, Cambridge MA
Hendry D, Doornik J (2014) Empirical Model Discovery and Theory Evaluation: Automatic Selection Methods in Econometrics. MIT Press, Cambridge MA
2014
-
[46]
Philosophia Mathematica 15:166--192
Hennig C (2007) Falsification of propensity models by statistical tests and the goodness-of-fit paradox. Philosophia Mathematica 15:166--192
2007
-
[47]
Foundations of Science 15:29--48
Hennig C (2010) Mathematical models and reality: a constructivist perspective. Foundations of Science 15:29--48
2010
-
[48]
Hoekstra R, Kiers H, Johnson A (2012) Are assumptions of well-known statistical techniques checked, and why (not)? Frontiers in Psychology 3:137
2012
-
[49]
In: Smelser NJ, Baltes PB (eds) International Encyclopedia of the Social and Behavioral Sciences, Pergamon, Oxford, pp 10673--10680
Hollander M, Sethuraman J (2001) Nonparametric statistics: Rank-based methods. In: Smelser NJ, Baltes PB (eds) International Encyclopedia of the Social and Behavioral Sciences, Pergamon, Oxford, pp 10673--10680
2001
-
[50]
Arthritis & Rheumatism 52:2495--2505
Holman AJ, Myers RR (2005) A randomized, double-blind, placebo-controlled trial of pramipexole, a dopamine agonist, in patients with fibromyalgia receiving concomitant medications. Arthritis & Rheumatism 52:2495--2505
2005
-
[51]
Statistical Research Memoirs 2:1--24
Hsu PL (1938) Contribution to the theory of ``student's'' t-test as applied to the problem of two samples. Statistical Research Memoirs 2:1--24
1938
-
[52]
American Educational Research Journal 6:515--527
Hsu TC, Feldt LS (1969) The effect of limitations on the number of criterion score values on the significance level of the F-test. American Educational Research Journal 6:515--527
1969
-
[53]
Statistics in Medicine 32:4540--4549, doi:10.1002/sim.5869
Kahan BC (2013) Bias in randomised factorial trials. Statistics in Medicine 32:4540--4549, doi:10.1002/sim.5869
2013 doi
-
[54]
PLoS Computational Biology 12:e1004961
Kass RE, Caffo BS, Davidian M, Meng XL, Yu B, Reid N (2016) Ten simple rules for effective statistical practice. PLoS Computational Biology 12:e1004961
2016
-
[55]
Review of Educational Research 68:350--386
Keselman HJ, Huberty CJ, Lix LM, Olejnik S, Cribbie RA, Donahue B, Kovalchuk RK, Lowman LL, Petoskey MD, Keselman JC, Levin JR (1998) Statistical practices of educational researchers: An analysis of their anova, manova, and ancova analyses. Review of Educational Research 68:350--386
1998
-
[56]
Keselman HJ, Othman AR, Wilcox RR (2013) Preliminary testing for normality: Is this a good practice? Journal of Modern Applied Statistical Methods 12:2--19
2013
-
[57]
Journal of Applied Science Research 2:296--300
Keskin S (2006) Comparison of several univariate normality tests regarding type i error rate and power of the test in simulation based small samples. Journal of Applied Science Research 2:296--300
2006
-
[58]
Journal of Econometrics 25:35--48
King ML, Giles DEA (1984) Autocorrelation pre-testing in the linear model: Estimation, testing and prediction. Journal of Econometrics 25:35--48
1984
-
[59]
Physiological Measurement 39:114010
Kokosinska D, Gieraltowski JJ, Zebrowski JJ, Orlowska-Baranowska E, Baranowski R (2018) Heart rate variability, multifractal multiscale patterns and their assessment criteria. Physiological Measurement 39:114010
2018
-
[60]
Annals of Statistics 44:907--927
Lee JD, Sun DL, Sun Y, Taylor JE (2016) Exact post-selection inference, with application to the lasso. Annals of Statistics 44:907--927
2016
-
[61]
Econometric Theory 21:21--59
Leeb H, P\" o tscher BM (2005) Model selection and inference: Fact and fiction. Econometric Theory 21:21--59
2005
-
[62]
Statistical Science 30:216--227
Leeb H, P\" o tscher BM, Ewald K (2015) On various confidence intervals post-model-selection. Statistical Science 30:216--227
2015
-
[63]
Springer, New York
Lehmann EL, Romano JP (2005) Testing Statistical Hypotheses. Springer, New York
2005
-
[64]
Levene H (1960) Robust tests for equality of variances. In: Olkin I, Ghurye SG, Hoeffding W, Madow WG, Mann HB (eds) Contributions to Probability and Statistics: Essays in Honor of Harold Hotelling, Stanford University Press, Redwood City CA, pp 278--292
1960
-
[65]
The American Statistician 44:131--136
Markowski CA, Markowski EP (1990) Conditions for the effectiveness of a preliminary test of variance. The American Statistician 44:131--136
1990
-
[66]
Methodology 5:131--136
Maydeu-Olivares A, Forero CA, Gallardo-Pujol D, Renom J (2009) Testing categorized bivariate normality with two-stage polychoric correlation estimates. Methodology 5:131--136
2009
-
[67]
Cambridge University Press, Cambridge
Mayo DG (2018) Statistical Inference as Severe Testing. Cambridge University Press, Cambridge
2018
-
[68]
Pakistan Journal of Information and Technology 2:135--139
Mendes M, Pala A (2003) Type i error rate and power of three normality tests. Pakistan Journal of Information and Technology 2:135--139
2003
-
[69]
Psychological Testing and Assessment Modeling 52:343--353
Moder K (2010) Alternatives to F-test in one way ANOVA in case of heterogeneity of variances (a simulation study). Psychological Testing and Assessment Modeling 52:343--353
2010
-
[70]
The American Statistician 46:19--21
Moser BK, Stevens GR (1992) Homogeneity of variance in the two-sample means test. The American Statistician 46:19--21
1992
-
[71]
Communications in Statistics-Theory and Methods 18:3963--3975
Moser BK, Stevens GR, Watts CL (1989) The two-sample t test versus satterthwaite's approximate f test. Communications in Statistics-Theory and Methods 18:3963--3975
1989
-
[72]
Technometrics 10:509--522
Neave HR, Granger CWJ (1968) A monte carlo study comparing various two sample tests for differences in means. Technometrics 10:509--522
1968
-
[73]
Neyman J (1952) Lectures and Conferences on Mathematical Statistics and Probability (2nd ed.). U.S. Department of Agriculture, Washington DC
1952
-
[74]
Journal of Family Medicine and Primary Care 5:24--33
Nour-Eldein H (2016) Statistical methods and errors in family medicine articles between 2010 and 2014-suez canal university, egypt: A cross-sectional study. Journal of Family Medicine and Primary Care 5:24--33
2016
-
[75]
Economics Letters 17:111--114
Ohtani K, Toyoda T (1985) Testing linear hypothesis on regression coefficients after a pre-test for disturbance variance. Economics Letters 17:111--114
1985
-
[76]
Philosophical Magazine 5:157--175
Pearson K (1900) On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Philosophical Magazine 5:157--175
1900
-
[77]
Supplement to the Journal of the Royal Statistical Society 4(1):119--130, ://www.jstor.org/stable/2984124
Pitman EJG (1937) Significance tests which may be applied to samples from any populations. Supplement to the Journal of the Royal Statistical Society 4(1):119--130, ://www.jstor.org/stable/2984124
1937
-
[78]
Communications in Statistics: Theory and Methods 11:109--126
Posten HO, Yeh HC, Owen DB (1982) Robustness of the two-sample t-test under violations of the homogeneity of variance assumptions. Communications in Statistics: Theory and Methods 11:109--126
1982
-
[79]
Psychology Science 46:175--208
Rasch D, Guiard V (2004) The robustness of parametric statistical methods. Psychology Science 46:175--208
2004
-
[80]
Statistical Papers 52:219--231
Rasch D, Kubinger KD, Moder K (2011) The two-sample t test: pre-testing its assumptions does not pay off. Statistical Papers 52:219--231
2011
-
[81]
Statistical Papers 53:531--547
Ravichandran J (2012) A review of preliminary test-based statistical methods for the benefit of six sigma quality practitioners. Statistical Papers 53:531--547
2012
-
[82]
Journal of Statistical Modeling and Analytics 2:21--33
Razali NM, Wah YB (2011) Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests. Journal of Statistical Modeling and Analytics 2:21--33
2011
-
[83]
British Journal of Mathematical and Statistical Psychology 64:410--426
Rochon J, Kieser M (2011) A closer look at the effect of preliminary goodness-of-fit testing for normality for the one-sample t-test. British Journal of Mathematical and Statistical Psychology 64:410--426
2011
-
[84]
BMC Medical Research Methodology 12:81--91
Rochon J, Gondan M, Kieser M (2012) To test or not to test: Preliminary assessment of normality when comparing two independent samples. BMC Medical Research Methodology 12:81--91
2012
-
[85]
Wiley-Interscience, Hoboken, NJ
Saleh AME (2006) Theory of preliminary test and Stein-type estimation with applications. Wiley-Interscience, Hoboken, NJ
2006
-
[86]
Statistics and Risk Modeling 1:455--478
Saleh AME, Sen PK (1983) Asymptotic properties of tests of hypothesis following a preliminary test. Statistics and Risk Modeling 1:455--478
1983
-
[87]
Biometrics Bulletin 2:110--114
Satterthwaite FE (1946) An approximate distribution of estimates of variance components. Biometrics Bulletin 2:110--114
1946
-
[88]
Journal of the American Statistical Association 65:1501--1508
Scheff\' e H (1970) Practical solutions of the behrens-fisher problem. Journal of the American Statistical Association 65:1501--1508
1970
-
[89]
Clinical and Experimental Dermatology 31:757--761
Schoder V, Himmelmann A, Wilhelm KP (2006 a ) Preliminary testing for normality: some statistical aspects of a common concept. Clinical and Experimental Dermatology 31:757--761
2006
-
[90]
Communications in Statistics - Theory and Methods 35:2275--2286
Schoder V, Himmelmann A, Wilhelm KP (2006 b ) Preliminary testing for normality: some statistical aspects of a common concept. Communications in Statistics - Theory and Methods 35:2275--2286
2006
-
[91]
Cambridge University Press, Cambridge
Spanos A (1999) Probability Theory and Statistical Inference: Econometric Modeling with Observational Data. Cambridge University Press, Cambridge
1999
-
[92]
Journal of Econometrics 158:204--220
Spanos A (2010) Akaike-type criteria and the reliability of inference: Model selection versus statistical model specification. Journal of Econometrics 158:204--220
2010
-
[93]
Journal of Economic Surveys 32:541--577
Spanos A (2018) Mis-specification testing in retrospect. Journal of Economic Surveys 32:541--577
2018
-
[94]
Journal of Scientometric Research 4:10--13
Sridharan K, Gowri S (2015) Reporting quality of statistics in indian journals: Analysis of articles over a period of two years. Journal of Scientometric Research 4:10--13
2015
-
[95]
The American Statistician 61:47--55
Strasak AM, Zaman Q, Marinell G, Pfeiffer KP, Ulmer H (2007a) The use of statistics in medical research: A comparison of the new england journal of medicine and nature medicine. The American Statistician 61:47--55
2007
-
[96]
Austrian Journal of Statistics 36:141--152
Strasak AM, Zaman Q, Marinell G, Pfeiffer KP, Ulmer H (2007b) The use of statistics in medical research: A comparison of wiener klinische wochenschrift and wiener medizinische wochenschrift. Austrian Journal of Statistics 36:141--152
2007
-
[97]
Journal of Econometrics 31:67--80
Toyoda T, Ohtani K (1986) Testing equality between sets of coefficients after a preliminary test for equality of disturbance variances in two linear regressions. Journal of Econometrics 31:67--80
1986
-
[98]
Journal of Statistical Planning and Inference 57:21--28
Tukey JW (1997) More honest foundations for data analysis. Journal of Statistical Planning and Inference 57:21--28
1997
-
[99]
The American Statistician 70(2):129--133, doi:10.1080/00031305.2016.1154108
Wasserstein RL, Lazar NA (2016) The asa statement on p-values: Context, process, and purpose. The American Statistician 70(2):129--133, doi:10.1080/00031305.2016.1154108
2016 arXiv
-
[100]
Biometrika 29:350--62
Welch BL (1938) The significance of the difference between two means when the population variances are unequal. Biometrika 29:350--62
1938
-
[101]
Biometrika 34:28--35
Welch BL (1947) The generalisation of student's problem when several different population variances are involved. Biometrika 34:28--35
1947
-
[102]
Psychology Science 49:2--12
Wiedermann W, Alexandrowicz R (2007) A plea for more general tests than those for location only: Further considerations on rasch & guiard's `the robustness of parametric statistical methods'. Psychology Science 49:2--12
2007
-
[103]
The Scientific World Journal 11:2106--2114
Wu S, Jin Z, Wei X, Gao Q, Lu J, Ma X, Wu C, He Q, Wu M, Wang R, Xu, He (2011) Misuse of statistical methods in 10 leading chinese medical journals in 1998 and 2008. The Scientific World Journal 11:2106--2114
2011
-
[104]
Journal of Construction Engineering and Management 145:04019049
Wu W, Hartless J, Tesei A, Gunji V, Ayer S, London J (2019) Design assessment in virtual and mixed reality environments: Comparison of novices and experts. Journal of Construction Engineering and Management 145:04019049
2019
-
[105]
Statistical Methodology 3:351--374
Zimmerman DW (2006) Two separate effects of variance heterogeneity on the validity and power of significance tests of location. Statistical Methodology 3:351--374
2006
-
[106]
British Journal of Mathematical and Statistical Psychology 64:388--409
Zimmerman DW (2011) A simple and effective decision rule for choosing a significance test to protect against non-normality. British Journal of Mathematical and Statistical Psychology 64:388--409
2011
-
[107]
British Journal of Mathematical and Statistical Psychology 67:1--29
Zimmerman DW (2014) Consequences of choosing samples in hypothesis testing to ensure homogeneity of variance. British Journal of Mathematical and Statistical Psychology 67:1--29
2014
-
[108]
Hennig, C. (2010). Mathematical models and reality: a constructivist perspective. Foundations of Science, 15, 29--48
2010
-
[109]
Hollander and J
M. Hollander and J. Sethuraman (2001)
2001
-
[110]
and Myers, Robin R
Holman, Andrew J. and Myers, Robin R. (2005) A randomized, double-blind, placebo-controlled trial of pramipexole, a dopamine agonist, in patients with fibromyalgia receiving concomitant medications. Arthritis & Rheumatism 52, 2495--2505
2005
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.