Pith. sign in

REVIEW 3 major objections 4 minor 110 references

Should we test the model assumptions before running a model-based test?

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that, under four explicit conditions, a procedure that tests model assumptions first and then chooses between two tests can have strictly higher power than either test applied unconditionally.

desk verdict A careful, honest survey plus a correct but simple lemma; the practical conclusion overreaches because the key independence relaxation is unproved, but it deserves serious refereeing. read the letter →

arxiv 1908.02218 v5 pith:FECCQTCX submitted 2019-08-06 stat.ME

classification stat.ME MSC 62F0362G10
keywords misspecificationtestinggoodness-of-fitcombinedproceduretwo-stagepreliminaryparadoxmodelassumptionsstatisticalpower
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether the common practice of testing a model's assumptions before running a model-based test is statistically sound. It argues that the literature's largely negative verdict is too pessimistic, and that a combined procedure—run a misspecification test, then apply the constrained test if it passes and an alternative unconstrained test if it fails—can be superior to running either test unconditionally. The key result, Lemma 1, shows this gain is possible whenever the misspecification test is informative and the two main tests are each best in one of two regimes, and when the misspecification test is independent of the main tests. The practical message is that model checking should target violations that matter for the main test, and should be evaluated in a mixture setting where both valid and violated assumptions can occur.

What carries the argument

The central object is the combined test $\Phi_C(z)=\Phi_{\mathrm{MC}}(z)$ if $\Phi_{\mathrm{MS}}(z)=0$, and $\Phi_{\mathrm{AU}}(z)$ otherwise. The proof machinery is a two-point mixture setup where the whole dataset is drawn from $P_\theta$ with probability $\lambda$ and from $Q$ with probability $1-\lambda$, making the powers of all three procedures linear in $\lambda$. Assumptions (I)–(III) say the MC test beats the AU test under $P_\theta$, the AU test beats the MC test under $Q$, and the misspecification test rejects the model more often under $Q$; assumption (IV) is independence of both main rejection events from the misspecification decision under both distributions. Lemma 1 combines these to show that at the crossing point $\lambda^*=\Delta_Q/(\Delta_\theta+\Delta_Q)$, the combined procedure's power exceeds both the MC and AU power by the positive term $\frac{\Delta_\theta\Delta_Q}{\Delta_\theta+\Delta_Q}(\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}})$.

What would settle it

Simulate a two-sample setup where the misspecification test and the main test are strongly dependent, for example the Shapiro-Wilk normality test and Welch's t-test on the same small normal sample, and compute the four conditional-probability deviations from independence in the paper's $\delta$ relaxation. Then check whether at $\lambda^*=\Delta_Q/(\Delta_\theta+\Delta_Q)$ the combined procedure's power still exceeds both unconditional powers; if it does not, condition (IV) is essential and the promised relaxation fails in that case.

Watch

Extended reading notes

Core claim

The central claim is that preliminary misspecification testing—checking model assumptions before a model-based test—can be beneficial, contrary to much of the existing literature. The authors formalize a two-regime setup: data come either from a distribution $P_\theta$ satisfying the model assumptions (where the constrained MC test has higher power) or from $Q$ violating them (where the unconstrained AU test has higher power), with $\lambda$ the probability of the first regime. Under assumptions (I)–(IV), Lemma 1 establishes that there is a $\lambda^*\in(0,1)$ such that the combined procedure has strictly higher power than both unconditional tests at that mixture. The advantage over the better of the two unconditional tests equals $\frac{\Delta_\theta\Delta_Q}{\Delta_\theta+\Delta_Q}(\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}})>0$, where $\Delta_\theta$ and $\Delta_Q$ are the power gaps and $\alpha_{\mathrm{MS}}$, $\alpha^*_{\mathrm{MS}}$ are the misspecification test's rejection probabilities under $P_\theta$ and $Q$. Section 7 broadens this into the claim that model checking is more useful than the literature suggests, provided conditions (a)–(d) hold.

Load-bearing premise

The load-bearing premise is the independence of the misspecification test from both main tests (assumption (IV)); the authors call it very restrictive and unrealistic in most situations, and their promised relaxation to approximate independence with a small $\delta$ is stated without proof.

Editorial extensions

If this is right

  • A combined procedure can strictly outperform both of its constituent tests in power for some mixture of valid and violated model assumptions, not just match the better one.
  • The conditions specify what makes model checking worthwhile: an informative misspecification test, approximate independence between the misspecification test and the main tests, and clear power advantages of both main tests in their own regimes.
  • Model checking should be aimed at detecting violations that are problematic for the main test, not at verifying that assumptions hold exactly.
  • Evaluating combined procedures only at $\lambda=0$ or $\lambda=1$, as most of the literature does, is too pessimistic; the mixture perspective can be more favorable.
  • In the two-sample normality example, the combined procedure performs close to the better of the t-test and the Wilcoxon-Mann-Whitney test, and can even exceed both in power for a skew-normal distribution, though its type 1 error can be mildly anti-conservative.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: A practical corollary the authors leave implicit is that the misspecification test's significance level could be tuned to maximize the positive advantage term, since $\alpha^*_{\mathrm{MS}}-\alpha_{\mathrm{MS}}$ enters linearly; higher misspecification-test levels may be better when the cost of failing to switch to the AU test is large.
  • Extension: The $\lambda$-mixture idea can be carried further: with a prior over a continuous family of violations, the same linearity argument would give conditions for the combined procedure to beat 'always use AU' or 'always use MC' in average power.
  • Extension: Because assumption (IV) is known to fail in common designs such as crossover trials, a quantitative bound on the power loss as the dependence parameter $\delta$ grows is needed; the paper promises but does not deliver this, so the practical message currently rests on a plausibility argument.
  • Extension: The framework suggests a diagnostic for practice: before adopting a check-then-test protocol, simulate the two regimes and the misspecification test's sensitivity to see whether the four conditions hold for the specific tests and plausible violation distributions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper asks whether model assumptions should be tested before running a model-based test, and argues that the literature is too pessimistic about such 'combined procedures.' It formalizes a protocol involving a misspecification (MS) test, a model-based constrained (MC) test, and an alternative unconstrained (AU) test, reviews the existing literature, and presents a new theoretical result. Lemma 1 states that under four assumptions—including independence of the MS rejection event from the MC and AU rejection events—there exists a mixture weight λ in (0,1) such that the combined procedure has strictly higher power than either the MC or the AU test. The paper also provides a simulation example (Table 1) and a deferred simulation summarized by Figure 1, and concludes with conditions (a)-(d) under which model checking is worthwhile.

Significance. The paper's main value is conceptual and synthetic: it frames the debate over preliminary model testing in a general way, identifies the needed ingredients for a successful combined procedure, and gives a simple algebraic result showing that under explicit assumptions a combined procedure can beat both unconditional tests. The lemma itself is correct under its stated assumptions and is a useful counterpoint to the uniformly negative conclusions in parts of the literature. However, the practical message depends on an unproved relaxation of the restrictive independence assumption (IV), and the empirical support is partial and partly deferred. If the relaxation can be proved or convincingly demonstrated by simulation, the paper would make a solid contribution; as it stands, the central claim is defensible but not fully established.

major comments (3)
  1. [Section 6, Lemma 1 and the paragraph following it] The practical conclusion of the paper rests on relaxing assumption (IV), the independence of RMC and RAU from RMS, which the authors themselves call 'very restrictive' and 'unrealistic in most situations.' The relaxation to approximate independence is asserted only in the sentence 'it can be relaxed ... at the price of a more tedious proof that we do not present here,' with no statement of the required bound on δ in terms of Δθ, ΔQ, αMS, and α*MS. Since the proof of Lemma 1 uses exact independence to replace conditional rejection probabilities by unconditional ones, any violation of (IV) enters directly into the sign of Pλ(RC) - Pλ(RMC). Without a quantified condition or a numerical demonstration that the power advantage survives realistic dependence, the central claim that model checking is more useful than the literature suggests is not established.
  2. [Section 6, Figure 1 and Example 1/Table 1] The claim that the range of λ for which the combined procedure is best is 'quite large' is supported only by simulations that are 'published elsewhere' and by Figure 1, which is described as a typical pattern without reporting sample sizes, numbers of replications, or other methodological details. The only numerical evidence in the paper, Table 1, shows the combined procedure to be anti-conservative for the normal, t3, and skew-normal cases (type 1 error probabilities 0.0512, 0.0515, and 0.0531 against the nominal 0.05) and clearly worse than both the t-test and the permutation test for the exponential distribution in terms of type 2 error (0.4849 versus 0.3389 for the t-test). Thus the empirical support for the practical message is partial, and the power advantage in Lemma 1 is not accompanied by a demonstration that the combined procedure respects the nominal level.
  3. [Section 6, Lemma 1 and Section 7] The positive result is an existence statement: it guarantees some λ in (0,1) with a strict power advantage, but it gives no expression for the length of the interval of such λ, and the proof shows only that the power difference at λ* is positive under exact independence. Even if the independence relaxation were proved, the practical recommendation in Section 7 would require the difference to be positive over a substantial range of λ, which is not established by the lemma. The authors acknowledge that simulations indicate a large range, but those simulations are deferred; as a mathematical result, Lemma 1 alone is too weak to support the broad conclusion that model checking is worthwhile.
minor comments (4)
  1. [Example 1, text preceding Table 1] There is a typo in the sentence 'The WMW test is is clearly superior to the t-test for the t3- and skew normal distribution'; the word 'is' is repeated.
  2. [Figure 1 caption] The caption would benefit from reporting the sample sizes and the number of simulation replicates; as it stands, the reader cannot evaluate the stability of the reported pattern.
  3. [Section 6, paragraph on the mixture setup] The distinction between the mixture over whole datasets and an ordinary mixture model over individual observations is clear, but the sentence 'the setup has a certain Bayesian flavor' could be expanded to explain how a frequentist should interpret the randomness in λ; at present the interpretability of Pλ as a sampling distribution is not fully discussed.
  4. [References] A few references are incomplete, for example Campbell (2019) is cited as an arXiv preprint with a DOI and no final publication details, and the entries for Campbell and Dean and for Rasch et al. lack page ranges in the reference list.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Lemma 1 is a conditional derivation under explicit assumptions; the unproved δ-relaxation of (IV) is a stated gap, not a circular step.

full rationale

The paper's central derivation is Lemma 1 (Section 6). Its proof is self-contained algebraic manipulation from assumptions (I)-(IV): it defines Δθ, ΔQ, αMS, α*MS, locates λ*=ΔQ/(Δθ+ΔQ), and shows Pλ*(RC)=Pλ*(RMC)+ΔθΔQ/(Δθ+ΔQ)(α*MS−αMS)=Pλ*(RAU)+the same positive term. Nowhere is the conclusion assumed; (I)-(III) are ordering and usefulness conditions, and (IV) is an independence condition, not a restatement of the target inequality. The paper explicitly flags (IV) as 'very restrictive' and 'unrealistic in most situations', and the promised relaxation to small δ is asserted without proof ('at the price of a more tedious proof that we do not present here'). This is a genuine correctness/generalisability gap that the practical message leans on, but it is not circularity: the exact-(IV) Lemma 1 is proven, and the gap is an unquantified continuity step, not an input disguised as an output. Self-citations (Hennig 2007, 2010; Gelman and Hennig 2017) supply terminology and framing ('misspecification paradox', constructivist model view) and are not load-bearing in the proof. The simulation example (Table 1) is external evidence with stated limitations (anti-conservativity; exponential case), not a fit-then-predict scheme. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no known result is relabelled as new. Score 0.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The central lemma has no fitted parameters: it is a conditional statement given assumptions (I)-(IV). The illustrative Example 2 uses user-chosen mixture weights λ = 1/2, 1/4, 1/4 that are not fitted to data. The main load-bearing assumptions are (I)-(IV), of which (IV) is the most fragile and is acknowledged as unrealistic; its relaxation is asserted without proof.

assumptions (8)
  • domain assumption Probability mixture form: a dataset comes entirely from Pθ with probability λ and entirely from Q with probability 1-λ.
    This two-distribution superpopulation is the novel setup of Section 6; it must be a reasonable description of the researcher's world for Lemma 1 to apply.
  • domain assumption (I) Δθ = Pθ(RMC) - Pθ(RAU) > 0: the constrained test has better power when the model holds.
    Required for combined procedure to be beneficial; if false, always using the AU test is better.
  • domain assumption (II) ΔQ = Q(RAU) - Q(RMC) > 0: the unconstrained test has better power when the model is violated.
    Required; if false, always using the MC test is better.
  • domain assumption (III) α*_MS = Q(RMS) > α_MS = Pθ(RMS): the misspecification test can distinguish the two regimes at least somewhat.
    Without this, the MS test carries no information to choose between MC and AU.
  • ad hoc to paper (IV) RMC and RAU are independent of RMS under both Pθ and Q.
    Acknowledged unrealistic; the paper says it can be relaxed to approximate independence but omits the proof.
  • domain assumption There is a well-defined extension M* of the main null hypothesis into the larger model M, with M* ∩ MΘ = MΘ0.
    Imported from Section 3; needed to give the unconstrained test a meaningful hypothesis.
  • domain assumption Level and power are the appropriate criteria for comparing procedures.
    The paper's recommendations are framed in terms of type 1 and type 2 error probabilities, which is a standard but non-neutral choice.
  • standard math Standard probability linearity for mixtures, and the algebra in the proof of Lemma 1.
    Uncontroversial background for computing Pλ probabilities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Should we test the model assumptions before running a model-based test?." pith.science (2026). https://pith.science/paper/FECCQTCX

@misc{pith2026190802218,
  author       = {Pith},
  title        = {Pith review of: Should we test the model assumptions before running a model-based test?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FECCQTCX}},
  note         = {Machine review of arXiv:1908.02218}
}
read the original abstract

Statistical methods are based on model assumptions, and it is statistical folklore that a method's model assumptions should be checked before applying it. This can be formally done by running one or more misspecification tests of model assumptions before running a method that requires these assumptions; here we focus on model-based tests. A combined test procedure can be defined by specifying a protocol in which first model assumptions are tested and then, conditionally on the outcome, a test is run that requires or does not require the tested assumptions. Although such an approach is often taken in practice, much of the literature that investigated this is surprisingly critical of it. Our aim is to explore conditions under which model checking is advisable or not advisable. For this, we review results regarding such "combined procedures" in the literature, we review and discuss controversial views on the role of model checking in statistics, and we present a general setup in which we can show that preliminary model checking is advantageous, which implies conditions for making model checking worthwhile.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

110 extracted references · 79 canonical work pages

  1. [1]

    Journal of Transportation Technologies 7:133--147

    Abdulhafedh A (2017) How to detect and remove temporal autocorrelation in vehicular crash data. Journal of Transportation Technologies 7:133--147

  2. [2]

    Journal of Statistical Planning and Inference 88:47--57

    Albers W, Boon PC, Kallenberg WC (2000a) The asymptotic behavior of tests for normal means based on a variance pre-test. Journal of Statistical Planning and Inference 88:47--57

  3. [3]

    Annals of Statistics 28:195--214

    Albers W, Boon PC, Kallenberg WC (2000b) Size and power of pretest procedures. Annals of Statistics 28:195--214

  4. [4]

    Journal of the American Statistical Association 65:1590--1596

    Arnold BC (1970) Hypothesis testing incorporating a preliminary test of significance. Journal of the American Statistical Association 65:1590--1596

  5. [5]

    Annals of Mathematical Statistics 27:1115--1122

    Bahadur R, Savage L (1956) The nonexistence of certain statistical procedures in nonparametric problems. Annals of Mathematical Statistics 27:1115--1122

  6. [6]

    Annals of Mathematical Statistics 15:190--204

    Bancroft TA (1944) On biases in estimation due to the use of preliminary tests of significance. Annals of Mathematical Statistics 15:190--204

  7. [7]

    Biometrics 20:427--442

    Bancroft TA (1964) Analysis and inference for incompletely specified models involving the use of preliminary test(s) of significance. Biometrics 20:427--442

  8. [8]

    International Statistical Review 45:117--127

    Bancroft TA, Han C (1977) Inference based on conditional specification: A note and a bibliography. International Statistical Review 45:117--127

Show all 110 references
  1. [9]

    Mathematical Proceedings of the Cambridge Philosophical Society 31:223--231

    Bartlett MS (1935) The effect of non-normality on the t distribution. Mathematical Proceedings of the Cambridge Philosophical Society 31:223--231

  2. [10]

    Annals of Statistics 41:802--837

    Berk R, Brown L, Buja A, Zhang K, Zhao L (2013) Berk, r., brown, l., buja, a., zhang, k., zhao, l. Annals of Statistics 41:802--837

  3. [11]

    International Journal of Approximate Reasoning 66:53--72

    Bickel DR (2015) Inference after checking multiple bayesian models for data conflict and applications to mitigating the influence of rejected priors. International Journal of Approximate Reasoning 66:53--72

  4. [12]

    arXiv:191100115 (stat), accepted by Methods in Ecology and Evolution doi:https://doi.org/10.1111/2041-210X.13559, ://arxiv.org/abs/1911.00115

    Campbell H (2019) The consequences of checking for zero-inflation and overdispersion in the analysis of count data. arXiv:191100115 (stat), accepted by Methods in Ecology and Evolution doi:https://doi.org/10.1111/2041-210X.13559, ://arxiv.org/abs/1911.00115

  5. [13]

    Statistics in Medicine 33(6):1042--1056, doi:https://doi.org/10.1002/sim.6021

    Campbell H, Dean C (2014) The consequences of proportional hazards based model selection. Statistics in Medicine 33(6):1042--1056, doi:https://doi.org/10.1002/sim.6021

  6. [14]

    Journal of the Royal Statistical Society, Series B 158:419--466

    Chatfield C (1995) Model uncertainty, data mining and statistical inference (with discussion). Journal of the Royal Statistical Society, Series B 158:419--466

  7. [15]

    Cambridge University Press, Cambridge

    Cox DR (2006) Principles of Statistical Inference. Cambridge University Press, Cambridge

  8. [16]

    Australian Journal of Statistics 22:143--153

    Cressie N (1980) Relaxing assumptions in the one-sample t-test. Australian Journal of Statistics 22:143--153

  9. [17]

    Chapman & Hall/CRC, Boca Raton FL

    Davies PL (2014) Data Analysis and Approximate Models. Chapman & Hall/CRC, Boca Raton FL

  10. [18]

    Annals of Statistics 16:1390--1420

    Donoho D (1988) One-sided inference about functionals of a density. Annals of Statistics 16:1390--1420

  11. [19]

    Wiley, New York

    Dowdy S, Wearden S, Chilko D (2004) Statistics for Research. Wiley, New York

  12. [20]

    Journal of the Royal Statistical Society, Series B 57:45--97

    Draper D (1995) Assessment and propagation of model uncertainty (with discussion). Journal of the Royal Statistical Society, Series B 57:45--97

  13. [21]

    Technometrics 18:1--9

    Easterling RG (1976) Goodness of fit and parameter estimation. Technometrics 18:1--9

  14. [22]

    Journal of Statistical Computation and Simulation 8:1--11

    Easterling RG, Anderson HE (1978) The effect of preliminary normality goodness of fit tests on subsequent inference. Journal of Statistical Computation and Simulation 8:1--11

  15. [23]

    Journal of the American Statistical Association 109:991--1007

    Efron B (2014) Estimation and accuracy after model selection. Journal of the American Statistical Association 109:991--1007

  16. [24]

    Journal of Statistical Computation and Simulation 76:803--816

    Farrell PJ, Rogers-Stewart K (2006) Comprehensive study of tests for normality and symmetry: extending the spiegelhalter test. Journal of Statistical Computation and Simulation 76:803--816

  17. [25]

    Statistics Surveys 4:1--39

    Fay MP, Proschan MA (2010) Wilcoxon-mann-whitney or t-test? on assumptions for hypothesis tests and multiple interpretations of decision rules. Statistics Surveys 4:1--39

  18. [26]

    Wiley, New York

    de Finetti B (1974) Theory of Probability. Wiley, New York

  19. [27]

    Econometrica: Journal of the Econometric Society 29:139--170

    Fisher FM (1961) On the cost of approximate specification in simultaneous equation estimation. Econometrica: Journal of the Econometric Society 29:139--170

  20. [28]

    Philosophical Transactions of the Royal Society of London A 22:309--368

    Fisher RA (1922) On the mathematical foundation of theoretical statistics. Philosophical Transactions of the Royal Society of London A 22:309--368

  21. [29]

    Statistics in Medicine 8:1421--1432

    Freeman P (1989) The performance of the two-stage analysis of two-treatment, two-period cross-over trials. Statistics in Medicine 8:1421--1432

  22. [30]

    BMC Dermatology 2:163--174

    Gambichler T, Bader A, Vojvodic M, Bechara FG, Sauermann K, Altmeyer P, Hoffmann K (2002) Impact of uva exposure on psychological parameters and circulating serotonin and melatonin. BMC Dermatology 2:163--174

  23. [31]

    Gamble C, Krishan A, Stocken D, Lewis S, Juszczak E, Doré C, Williamson PR, Altman DG, Montgomery A, Lim P, Berlin J, Senn S, Day S, Barbachano Y, Loder E (2017) Guidelines for the Content of Statistical Analysis Plans in Clinical Trials . JAMA 318(23):2337--2343, doi:10.1001/...

  24. [32]

    Communications in Statistics - Simulation and Computation 10:163--174

    Gans DJ (1981) Use of a preliminary test in comparing two sample means. Communications in Statistics - Simulation and Computation 10:163--174

  25. [33]

    Journal of the Royal Statistical Society, Series A 180:967--1033

    Gelman A, Hennig C (2017) Beyond subjective and objective in statistics (with discussion). Journal of the Royal Statistical Society, Series A 180:967--1033

  26. [34]

    The American Statistician 102:460--465

    Gelman A, Loken E (2014) The statistical crisis in science. The American Statistician 102:460--465

  27. [35]

    British Journal of Mathematical and Statistical Psychology 66:8--38

    Gelman A, Shalizi CR (2013) Philosophy and the practice of bayesian statistics. British Journal of Mathematical and Statistical Psychology 66:8--38

  28. [36]

    Journal of Economic Surveys 7:145--197

    Giles DEA, Giles JA (1993) Pre-test estimation and testing in econometrics: Recent developments. Journal of Economic Surveys 7:145--197

  29. [37]

    The Lagrange Multiplier principle and other applications

    Godfrey LG (1988) Misspecification tests in econometrics. The Lagrange Multiplier principle and other applications. Cambridge University Press, Cambridge

  30. [38]

    Journal of Statistical Planning and Inference 49:241--260

    Godfrey LG (1996) Misspecification tests and their uses in econometrics. Journal of Statistical Planning and Inference 49:241--260

  31. [39]

    Biometrics 21:469--480 (Corrigendum in Biometrics, 30, 727, 1974)

    Grizzle JE (1967) The two-period change-over design and its use in clinical trials. Biometrics 21:469--480 (Corrigendum in Biometrics, 30, 727, 1974)

  32. [40]

    Journal of the Indian Statistical Association 7:26--29

    Gupta VP, Srivastava VK (1993) Upper bound for the size of a test procedure using preliminary tests of significance. Journal of the Indian Statistical Association 7:26--29

  33. [41]

    Biometrika 49:403--417

    Gurland J, McCullough R (1962) Testing equality of means after a preliminary test of equality of variances. Biometrika 49:403--417

  34. [42]

    Wiley, New York

    Hampel FR, Ronchetti EM, Rousseeuw PJ, Stahel WA (1986) Robust Statistics. Wiley, New York

  35. [43]

    Frontiers in Neuroscience 14:687

    Hasler G, Suker S, Schoretsanitis G, Mihov Y (2020) Sustained improvement of negative self-schema after a single ketamine infusion: An open-label study. Frontiers in Neuroscience 14:687

  36. [44]

    Journal of the American Statistical Association 85:446--452

    He X, Simpson DG, Portnoy SL (1990) Breakdown robustness of tests. Journal of the American Statistical Association 85:446--452

  37. [45]

    MIT Press, Cambridge MA

    Hendry D, Doornik J (2014) Empirical Model Discovery and Theory Evaluation: Automatic Selection Methods in Econometrics. MIT Press, Cambridge MA

  38. [46]

    Philosophia Mathematica 15:166--192

    Hennig C (2007) Falsification of propensity models by statistical tests and the goodness-of-fit paradox. Philosophia Mathematica 15:166--192

  39. [47]

    Foundations of Science 15:29--48

    Hennig C (2010) Mathematical models and reality: a constructivist perspective. Foundations of Science 15:29--48

  40. [48]

    Hoekstra R, Kiers H, Johnson A (2012) Are assumptions of well-known statistical techniques checked, and why (not)? Frontiers in Psychology 3:137

  41. [49]

    In: Smelser NJ, Baltes PB (eds) International Encyclopedia of the Social and Behavioral Sciences, Pergamon, Oxford, pp 10673--10680

    Hollander M, Sethuraman J (2001) Nonparametric statistics: Rank-based methods. In: Smelser NJ, Baltes PB (eds) International Encyclopedia of the Social and Behavioral Sciences, Pergamon, Oxford, pp 10673--10680

  42. [50]

    Arthritis & Rheumatism 52:2495--2505

    Holman AJ, Myers RR (2005) A randomized, double-blind, placebo-controlled trial of pramipexole, a dopamine agonist, in patients with fibromyalgia receiving concomitant medications. Arthritis & Rheumatism 52:2495--2505

  43. [51]

    Statistical Research Memoirs 2:1--24

    Hsu PL (1938) Contribution to the theory of ``student's'' t-test as applied to the problem of two samples. Statistical Research Memoirs 2:1--24

  44. [52]

    American Educational Research Journal 6:515--527

    Hsu TC, Feldt LS (1969) The effect of limitations on the number of criterion score values on the significance level of the F-test. American Educational Research Journal 6:515--527

  45. [53]

    Statistics in Medicine 32:4540--4549, doi:10.1002/sim.5869

    Kahan BC (2013) Bias in randomised factorial trials. Statistics in Medicine 32:4540--4549, doi:10.1002/sim.5869

  46. [54]

    PLoS Computational Biology 12:e1004961

    Kass RE, Caffo BS, Davidian M, Meng XL, Yu B, Reid N (2016) Ten simple rules for effective statistical practice. PLoS Computational Biology 12:e1004961

  47. [55]

    Review of Educational Research 68:350--386

    Keselman HJ, Huberty CJ, Lix LM, Olejnik S, Cribbie RA, Donahue B, Kovalchuk RK, Lowman LL, Petoskey MD, Keselman JC, Levin JR (1998) Statistical practices of educational researchers: An analysis of their anova, manova, and ancova analyses. Review of Educational Research 68:350--386

  48. [56]

    Keselman HJ, Othman AR, Wilcox RR (2013) Preliminary testing for normality: Is this a good practice? Journal of Modern Applied Statistical Methods 12:2--19

  49. [57]

    Journal of Applied Science Research 2:296--300

    Keskin S (2006) Comparison of several univariate normality tests regarding type i error rate and power of the test in simulation based small samples. Journal of Applied Science Research 2:296--300

  50. [58]

    Journal of Econometrics 25:35--48

    King ML, Giles DEA (1984) Autocorrelation pre-testing in the linear model: Estimation, testing and prediction. Journal of Econometrics 25:35--48

  51. [59]

    Physiological Measurement 39:114010

    Kokosinska D, Gieraltowski JJ, Zebrowski JJ, Orlowska-Baranowska E, Baranowski R (2018) Heart rate variability, multifractal multiscale patterns and their assessment criteria. Physiological Measurement 39:114010

  52. [60]

    Annals of Statistics 44:907--927

    Lee JD, Sun DL, Sun Y, Taylor JE (2016) Exact post-selection inference, with application to the lasso. Annals of Statistics 44:907--927

  53. [61]

    Econometric Theory 21:21--59

    Leeb H, P\" o tscher BM (2005) Model selection and inference: Fact and fiction. Econometric Theory 21:21--59

  54. [62]

    Statistical Science 30:216--227

    Leeb H, P\" o tscher BM, Ewald K (2015) On various confidence intervals post-model-selection. Statistical Science 30:216--227

  55. [63]

    Springer, New York

    Lehmann EL, Romano JP (2005) Testing Statistical Hypotheses. Springer, New York

  56. [64]

    Levene H (1960) Robust tests for equality of variances. In: Olkin I, Ghurye SG, Hoeffding W, Madow WG, Mann HB (eds) Contributions to Probability and Statistics: Essays in Honor of Harold Hotelling, Stanford University Press, Redwood City CA, pp 278--292

  57. [65]

    The American Statistician 44:131--136

    Markowski CA, Markowski EP (1990) Conditions for the effectiveness of a preliminary test of variance. The American Statistician 44:131--136

  58. [66]

    Methodology 5:131--136

    Maydeu-Olivares A, Forero CA, Gallardo-Pujol D, Renom J (2009) Testing categorized bivariate normality with two-stage polychoric correlation estimates. Methodology 5:131--136

  59. [67]

    Cambridge University Press, Cambridge

    Mayo DG (2018) Statistical Inference as Severe Testing. Cambridge University Press, Cambridge

  60. [68]

    Pakistan Journal of Information and Technology 2:135--139

    Mendes M, Pala A (2003) Type i error rate and power of three normality tests. Pakistan Journal of Information and Technology 2:135--139

  61. [69]

    Psychological Testing and Assessment Modeling 52:343--353

    Moder K (2010) Alternatives to F-test in one way ANOVA in case of heterogeneity of variances (a simulation study). Psychological Testing and Assessment Modeling 52:343--353

  62. [70]

    The American Statistician 46:19--21

    Moser BK, Stevens GR (1992) Homogeneity of variance in the two-sample means test. The American Statistician 46:19--21

  63. [71]

    Communications in Statistics-Theory and Methods 18:3963--3975

    Moser BK, Stevens GR, Watts CL (1989) The two-sample t test versus satterthwaite's approximate f test. Communications in Statistics-Theory and Methods 18:3963--3975

  64. [72]

    Technometrics 10:509--522

    Neave HR, Granger CWJ (1968) A monte carlo study comparing various two sample tests for differences in means. Technometrics 10:509--522

  65. [73]

    Neyman J (1952) Lectures and Conferences on Mathematical Statistics and Probability (2nd ed.). U.S. Department of Agriculture, Washington DC

  66. [74]

    Journal of Family Medicine and Primary Care 5:24--33

    Nour-Eldein H (2016) Statistical methods and errors in family medicine articles between 2010 and 2014-suez canal university, egypt: A cross-sectional study. Journal of Family Medicine and Primary Care 5:24--33

  67. [75]

    Economics Letters 17:111--114

    Ohtani K, Toyoda T (1985) Testing linear hypothesis on regression coefficients after a pre-test for disturbance variance. Economics Letters 17:111--114

  68. [76]

    Philosophical Magazine 5:157--175

    Pearson K (1900) On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Philosophical Magazine 5:157--175

  69. [77]

    Supplement to the Journal of the Royal Statistical Society 4(1):119--130, ://www.jstor.org/stable/2984124

    Pitman EJG (1937) Significance tests which may be applied to samples from any populations. Supplement to the Journal of the Royal Statistical Society 4(1):119--130, ://www.jstor.org/stable/2984124

  70. [78]

    Communications in Statistics: Theory and Methods 11:109--126

    Posten HO, Yeh HC, Owen DB (1982) Robustness of the two-sample t-test under violations of the homogeneity of variance assumptions. Communications in Statistics: Theory and Methods 11:109--126

  71. [79]

    Psychology Science 46:175--208

    Rasch D, Guiard V (2004) The robustness of parametric statistical methods. Psychology Science 46:175--208

  72. [80]

    Statistical Papers 52:219--231

    Rasch D, Kubinger KD, Moder K (2011) The two-sample t test: pre-testing its assumptions does not pay off. Statistical Papers 52:219--231

  73. [81]

    Statistical Papers 53:531--547

    Ravichandran J (2012) A review of preliminary test-based statistical methods for the benefit of six sigma quality practitioners. Statistical Papers 53:531--547

  74. [82]

    Journal of Statistical Modeling and Analytics 2:21--33

    Razali NM, Wah YB (2011) Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests. Journal of Statistical Modeling and Analytics 2:21--33

  75. [83]

    British Journal of Mathematical and Statistical Psychology 64:410--426

    Rochon J, Kieser M (2011) A closer look at the effect of preliminary goodness-of-fit testing for normality for the one-sample t-test. British Journal of Mathematical and Statistical Psychology 64:410--426

  76. [84]

    BMC Medical Research Methodology 12:81--91

    Rochon J, Gondan M, Kieser M (2012) To test or not to test: Preliminary assessment of normality when comparing two independent samples. BMC Medical Research Methodology 12:81--91

  77. [85]

    Wiley-Interscience, Hoboken, NJ

    Saleh AME (2006) Theory of preliminary test and Stein-type estimation with applications. Wiley-Interscience, Hoboken, NJ

  78. [86]

    Statistics and Risk Modeling 1:455--478

    Saleh AME, Sen PK (1983) Asymptotic properties of tests of hypothesis following a preliminary test. Statistics and Risk Modeling 1:455--478

  79. [87]

    Biometrics Bulletin 2:110--114

    Satterthwaite FE (1946) An approximate distribution of estimates of variance components. Biometrics Bulletin 2:110--114

  80. [88]

    Journal of the American Statistical Association 65:1501--1508

    Scheff\' e H (1970) Practical solutions of the behrens-fisher problem. Journal of the American Statistical Association 65:1501--1508

  81. [89]

    Clinical and Experimental Dermatology 31:757--761

    Schoder V, Himmelmann A, Wilhelm KP (2006 a ) Preliminary testing for normality: some statistical aspects of a common concept. Clinical and Experimental Dermatology 31:757--761

  82. [90]

    Communications in Statistics - Theory and Methods 35:2275--2286

    Schoder V, Himmelmann A, Wilhelm KP (2006 b ) Preliminary testing for normality: some statistical aspects of a common concept. Communications in Statistics - Theory and Methods 35:2275--2286

  83. [91]

    Cambridge University Press, Cambridge

    Spanos A (1999) Probability Theory and Statistical Inference: Econometric Modeling with Observational Data. Cambridge University Press, Cambridge

  84. [92]

    Journal of Econometrics 158:204--220

    Spanos A (2010) Akaike-type criteria and the reliability of inference: Model selection versus statistical model specification. Journal of Econometrics 158:204--220

  85. [93]

    Journal of Economic Surveys 32:541--577

    Spanos A (2018) Mis-specification testing in retrospect. Journal of Economic Surveys 32:541--577

  86. [94]

    Journal of Scientometric Research 4:10--13

    Sridharan K, Gowri S (2015) Reporting quality of statistics in indian journals: Analysis of articles over a period of two years. Journal of Scientometric Research 4:10--13

  87. [95]

    The American Statistician 61:47--55

    Strasak AM, Zaman Q, Marinell G, Pfeiffer KP, Ulmer H (2007a) The use of statistics in medical research: A comparison of the new england journal of medicine and nature medicine. The American Statistician 61:47--55

  88. [96]

    Austrian Journal of Statistics 36:141--152

    Strasak AM, Zaman Q, Marinell G, Pfeiffer KP, Ulmer H (2007b) The use of statistics in medical research: A comparison of wiener klinische wochenschrift and wiener medizinische wochenschrift. Austrian Journal of Statistics 36:141--152

  89. [97]

    Journal of Econometrics 31:67--80

    Toyoda T, Ohtani K (1986) Testing equality between sets of coefficients after a preliminary test for equality of disturbance variances in two linear regressions. Journal of Econometrics 31:67--80

  90. [98]

    Journal of Statistical Planning and Inference 57:21--28

    Tukey JW (1997) More honest foundations for data analysis. Journal of Statistical Planning and Inference 57:21--28

  91. [99]

    The American Statistician 70(2):129--133, doi:10.1080/00031305.2016.1154108

    Wasserstein RL, Lazar NA (2016) The asa statement on p-values: Context, process, and purpose. The American Statistician 70(2):129--133, doi:10.1080/00031305.2016.1154108

  92. [100]

    Biometrika 29:350--62

    Welch BL (1938) The significance of the difference between two means when the population variances are unequal. Biometrika 29:350--62

  93. [101]

    Biometrika 34:28--35

    Welch BL (1947) The generalisation of student's problem when several different population variances are involved. Biometrika 34:28--35

  94. [102]

    Psychology Science 49:2--12

    Wiedermann W, Alexandrowicz R (2007) A plea for more general tests than those for location only: Further considerations on rasch & guiard's `the robustness of parametric statistical methods'. Psychology Science 49:2--12

  95. [103]

    The Scientific World Journal 11:2106--2114

    Wu S, Jin Z, Wei X, Gao Q, Lu J, Ma X, Wu C, He Q, Wu M, Wang R, Xu, He (2011) Misuse of statistical methods in 10 leading chinese medical journals in 1998 and 2008. The Scientific World Journal 11:2106--2114

  96. [104]

    Journal of Construction Engineering and Management 145:04019049

    Wu W, Hartless J, Tesei A, Gunji V, Ayer S, London J (2019) Design assessment in virtual and mixed reality environments: Comparison of novices and experts. Journal of Construction Engineering and Management 145:04019049

  97. [105]

    Statistical Methodology 3:351--374

    Zimmerman DW (2006) Two separate effects of variance heterogeneity on the validity and power of significance tests of location. Statistical Methodology 3:351--374

  98. [106]

    British Journal of Mathematical and Statistical Psychology 64:388--409

    Zimmerman DW (2011) A simple and effective decision rule for choosing a significance test to protect against non-normality. British Journal of Mathematical and Statistical Psychology 64:388--409

  99. [107]

    British Journal of Mathematical and Statistical Psychology 67:1--29

    Zimmerman DW (2014) Consequences of choosing samples in hypothesis testing to ensure homogeneity of variance. British Journal of Mathematical and Statistical Psychology 67:1--29

  100. [108]

    Hennig, C. (2010). Mathematical models and reality: a constructivist perspective. Foundations of Science, 15, 29--48

  101. [109]

    Hollander and J

    M. Hollander and J. Sethuraman (2001)

  102. [110]

    and Myers, Robin R

    Holman, Andrew J. and Myers, Robin R. (2005) A randomized, double-blind, placebo-controlled trial of pramipexole, a dopamine agonist, in patients with fibromyalgia receiving concomitant medications. Arthritis & Rheumatism 52, 2495--2505

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.