{"id":"d3a1ac80-0702-41a6-851a-a7cdb7e8d3e5","arxiv_id":"1908.02218","paper_version":5,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-stage procedure that checks model assumptions before choosing a test can beat always using one test, provided the check has some power and is nearly independent of both candidate tests.","lead":"This paper asks whether researchers should test statistical model assumptions before running a test that relies on them, and reviews the surprisingly negative literature on such two-stage procedures. It then proves that, under specific conditions, a combined procedure can be strictly more powerful than always using either the constrained or the unconstrained test.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1 is correct under exact (IV), but the paper's practical message depends on the unproved approximate-independence relaxation of (IV), which is asserted without proof or quantification.","rationale":"The reader's weakest_assumption is exactly the same load-bearing concern: Assumption (IV) and its unproved relaxation. I have checked the proof of Lemma 1 in Section 6 and it is algebraically correct; the independence assumption is used only in the equality Pθ(RC) = αMS Pθ(RAU) + (1-αMS) Pθ(RMC), and likewise for Q. The closed-form excess over both unconditional tests is (Δθ ΔQ / (Δθ + ΔQ))(α*MS - αMS) > 0. The paper's own Section 6 acknowledges that (IV) is unrealistic, and the claimed δ relaxation is asserted without proof. This is a genuine gap in the central argument: the positive result as stated applies only under exact independence, which the authors themselves call unrealistic, while the practical conclusion in Section 7 and the conditions (a)-(d) rely on approximate independence. The paper earns credit for transparency: it flags the limitation, provides a simulation-based Figure 1 (though the details are deferred to another publication), and frames the result as giving 'an idea of the required ingredients' rather than a universal proof. Given that the central Lemma is correct as stated and the flaw is a missing proof of a stated relaxation, the verdict CONDITIONAL is appropriate: accept the mathematical result, but condition the broader practical claim on either a proof of the δ relaxation or a numerical demonstration that the excess remains positive under realistic dependence. I do not see a basis for REJECT, because the authors' interpretive claim is explicitly conditional on (a)-(d) and the paper is honest about the restrictiveness. The concern is not about novelty or internal inconsistency of the proved result, but about the gap between what is proved and what is claimed for practice.","tokens_in":25348,"tokens_out":1738,"duration_ms":16478,"concrete_test":"Prove or numerically falsify the claimed relaxation of (IV): compute the exact or simulated quantity |Pθ(RMC|RMS) - Pθ(RMC|Rc_MS)| (and the analogous three quantities) for a concrete combined procedure—e.g., Shapiro-Wilk MS test, Welch t-test MC, WMW AU, under normal and t3 alternatives with the Section 6 simulation setup—and then directly verify whether the inequality Pλ(RC) > max(Pλ(RMC), Pλ(RAU)) still holds at λ* = ΔQ/(Δθ+ΔQ) for the observed conditional dependence. If the inequality fails for a realistically small δ, the claim that the relaxation is 'at the price of a more tedious proof' needs revision.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's practical claim that model checking is more useful than the literature suggests rests on Lemma 1. Lemma 1 itself is correct under (I)-(IV), including exact independence. However, the authors explicitly call (IV) 'very restrictive' and 'unrealistic in most situations,' then state that it can be relaxed to approximate independence with small δ 'at the price of a more tedious proof that we do not present here' (Section 6). This is an unproved assertion in the central argument. The sign of the Lemma's advantage is Pλ*(RC) - Pλ*(RMC) = [Δθ ΔQ / (Δθ + ΔQ)] (α*MS - αMS), which is exactly zero if conditional dependence is large enough to reverse the ordering of the conditional rejection probabilities. The envelope of δ that preserves the inequality is not derived, and no simulation or numerical check is provided outside Figure 1, whose supporting simulations are deferred to another publication. In the two-sample normality example, Table 1 actually shows the combined procedure can be anti-conservative, and for the exponential distribution the combined procedure performs poorly. Thus the practical lesson depends on a relaxation that is neither proved nor quantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether model assumptions should be tested before running a model-based test, and argues that the literature is too pessimistic about such 'combined procedures.' It formalizes a protocol involving a misspecification (MS) test, a model-based constrained (MC) test, and an alternative unconstrained (AU) test, reviews the existing literature, and presents a new theoretical result. Lemma 1 states that under four assumptions—including independence of the MS rejection event from the MC and AU rejection events—there exists a mixture weight λ in (0,1) such that the combined procedure has strictly higher power than either the MC or the AU test. The paper also provides a simulation example (Table 1) and a deferred simulation summarized by Figure 1, and concludes with conditions (a)-(d) under which model checking is worthwhile.","tokens_in":25601,"tokens_out":4937,"duration_ms":59131,"significance":"The paper's main value is conceptual and synthetic: it frames the debate over preliminary model testing in a general way, identifies the needed ingredients for a successful combined procedure, and gives a simple algebraic result showing that under explicit assumptions a combined procedure can beat both unconditional tests. The lemma itself is correct under its stated assumptions and is a useful counterpoint to the uniformly negative conclusions in parts of the literature. However, the practical message depends on an unproved relaxation of the restrictive independence assumption (IV), and the empirical support is partial and partly deferred. If the relaxation can be proved or convincingly demonstrated by simulation, the paper would make a solid contribution; as it stands, the central claim is defensible but not fully established.","major_comments":[{"comment":"The practical conclusion of the paper rests on relaxing assumption (IV), the independence of RMC and RAU from RMS, which the authors themselves call 'very restrictive' and 'unrealistic in most situations.' The relaxation to approximate independence is asserted only in the sentence 'it can be relaxed ... at the price of a more tedious proof that we do not present here,' with no statement of the required bound on δ in terms of Δθ, ΔQ, αMS, and α*MS. Since the proof of Lemma 1 uses exact independence to replace conditional rejection probabilities by unconditional ones, any violation of (IV) enters directly into the sign of Pλ(RC) - Pλ(RMC). Without a quantified condition or a numerical demonstration that the power advantage survives realistic dependence, the central claim that model checking is more useful than the literature suggests is not established.","section":"Section 6, Lemma 1 and the paragraph following it"},{"comment":"The claim that the range of λ for which the combined procedure is best is 'quite large' is supported only by simulations that are 'published elsewhere' and by Figure 1, which is described as a typical pattern without reporting sample sizes, numbers of replications, or other methodological details. The only numerical evidence in the paper, Table 1, shows the combined procedure to be anti-conservative for the normal, t3, and skew-normal cases (type 1 error probabilities 0.0512, 0.0515, and 0.0531 against the nominal 0.05) and clearly worse than both the t-test and the permutation test for the exponential distribution in terms of type 2 error (0.4849 versus 0.3389 for the t-test). Thus the empirical support for the practical message is partial, and the power advantage in Lemma 1 is not accompanied by a demonstration that the combined procedure respects the nominal level.","section":"Section 6, Figure 1 and Example 1/Table 1"},{"comment":"The positive result is an existence statement: it guarantees some λ in (0,1) with a strict power advantage, but it gives no expression for the length of the interval of such λ, and the proof shows only that the power difference at λ* is positive under exact independence. Even if the independence relaxation were proved, the practical recommendation in Section 7 would require the difference to be positive over a substantial range of λ, which is not established by the lemma. The authors acknowledge that simulations indicate a large range, but those simulations are deferred; as a mathematical result, Lemma 1 alone is too weak to support the broad conclusion that model checking is worthwhile.","section":"Section 6, Lemma 1 and Section 7"}],"minor_comments":[{"comment":"There is a typo in the sentence 'The WMW test is is clearly superior to the t-test for the t3- and skew normal distribution'; the word 'is' is repeated.","section":"Example 1, text preceding Table 1"},{"comment":"The caption would benefit from reporting the sample sizes and the number of simulation replicates; as it stands, the reader cannot evaluate the stability of the reported pattern.","section":"Figure 1 caption"},{"comment":"The distinction between the mixture over whole datasets and an ordinary mixture model over individual observations is clear, but the sentence 'the setup has a certain Bayesian flavor' could be expanded to explain how a frequentist should interpret the randomness in λ; at present the interpretability of Pλ as a sampling distribution is not fully discussed.","section":"Section 6, paragraph on the mixture setup"},{"comment":"A few references are incomplete, for example Campbell (2019) is cited as an arXiv preprint with a DOI and no final publication details, and the entries for Campbell and Dean and for Rasch et al. lack page ranges in the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a review-plus-simple-theorem contribution. Its novel mathematical content is modest and is correct only under an assumption the authors admit is unrealistic. The missing proof or quantitative treatment of the approximate-independence relaxation is the key gap. If the authors can provide that proof, or failing that, a careful simulation study demonstrating robustness of the conclusion to violations of (IV), the paper would be publishable. Without this, the manuscript's central practical claim is not sufficiently supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a careful, honest paper, but the headline claim is bigger than the proof. The genuinely new piece is a small lemma: if you are mixing two data-generating regimes, one where the constrained test is better and one where the unconstrained test is better, then under exact independence between the misspecification test and the two main tests, the combined procedure beats both unconditional tests for some mixture proportion. The lemma is correct; the algebra is simple and the authors openly call it 'not in itself a particularly strong result.' What matters is that they use it to argue that model checking is more useful than the literature suggests, and that inference depends on an unproved relaxation of the independence assumption.\n\nCredit where due: the survey is the real work here. Sections 4 and 5 bring together the variance-pooling, normality-testing, regression, and crossover-trial literatures, and the authors are fair to results that disagree with them. The misspecification paradox discussion is sensible, and the list of conditions (a)-(d) in Section 7 is a useful checklist for practitioners. Table 1 is informative, and they report the anti-conservative type 1 errors and the bad exponential case rather than hiding them. The self-citations are used for framing and terminology, not to prop up the proof. No circularity problem.\n\nSoft spots, in proportion. The biggest is assumption (IV), which the authors themselves call 'very restrictive' and 'unrealistic in most situations.' They say it can be relaxed to approximate independence if some deviations are smaller than delta, but they do not prove that, and they do not quantify delta. The sign of the advantage in Lemma 1 flips if dependence between the MS test and the main tests is strong enough, so this is not a cosmetic gap. Also, the simulation evidence behind Figure 1 is deferred to another publication, and the one worked example here is mixed: the combined procedure is good in the skew-normal case but anti-conservative, and poor for the exponential distribution. These gaps do not invalidate the lemma as stated, but they should cap how strongly the practical conclusion is phrased.\n\nI would send this to a serious referee. It is a legitimate contribution to a long-running debate, and the review portion alone has value. The right revision asks for either a proof of the approximate-independence relaxation, a quantitative bound on delta, or a reframed conclusion that explicitly conditions on (IV). With that, it would be a solid reference. I would likely cite it for the survey and for the lemma, and it would make a good reading-group discussion.\n\nRecommended action: engage with it, conditionally.","headline":"A careful, honest survey plus a correct but simple lemma; the practical conclusion overreaches because the key independence relaxation is unproved, but it deserves serious refereeing.","tokens_in":26085,"tokens_out":3147,"would_cite":true,"duration_ms":35777,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that, under four explicit conditions, a procedure that tests model assumptions first and then chooses between two tests can have strictly higher power than either test applied unconditionally.","keywords":["misspecification testing","goodness-of-fit testing","combined procedure","two-stage testing","preliminary testing","misspecification paradox","model assumptions","statistical power"],"falsifier":"Simulate a two-sample setup where the misspecification test and the main test are strongly dependent, for example the Shapiro-Wilk normality test and Welch's t-test on the same small normal sample, and compute the four conditional-probability deviations from independence in the paper's $\\delta$ relaxation. Then check whether at $\\lambda^*=\\Delta_Q/(\\Delta_\\theta+\\Delta_Q)$ the combined procedure's power still exceeds both unconditional powers; if it does not, condition (IV) is essential and the promised relaxation fails in that case.","tokens_in":25160,"feed_emoji":"📊","tokens_out":6372,"duration_ms":56697,"temperature":0.7,"pith_summary":"The paper asks whether the common practice of testing a model's assumptions before running a model-based test is statistically sound. It argues that the literature's largely negative verdict is too pessimistic, and that a combined procedure—run a misspecification test, then apply the constrained test if it passes and an alternative unconstrained test if it fails—can be superior to running either test unconditionally. The key result, Lemma 1, shows this gain is possible whenever the misspecification test is informative and the two main tests are each best in one of two regimes, and when the misspecification test is independent of the main tests. The practical message is that model checking should target violations that matter for the main test, and should be evaluated in a mixture setting where both valid and violated assumptions can occur.","feed_headline":"Pre-testing assumptions can beat both tests used alone","feed_subtitle":"Under four explicit conditions—notably independence—a check-then-test procedure strictly beats its two components in power.","key_machinery":"The central object is the combined test $\\Phi_C(z)=\\Phi_{\\mathrm{MC}}(z)$ if $\\Phi_{\\mathrm{MS}}(z)=0$, and $\\Phi_{\\mathrm{AU}}(z)$ otherwise. The proof machinery is a two-point mixture setup where the whole dataset is drawn from $P_\\theta$ with probability $\\lambda$ and from $Q$ with probability $1-\\lambda$, making the powers of all three procedures linear in $\\lambda$. Assumptions (I)–(III) say the MC test beats the AU test under $P_\\theta$, the AU test beats the MC test under $Q$, and the misspecification test rejects the model more often under $Q$; assumption (IV) is independence of both main rejection events from the misspecification decision under both distributions. Lemma 1 combines these to show that at the crossing point $\\lambda^*=\\Delta_Q/(\\Delta_\\theta+\\Delta_Q)$, the combined procedure's power exceeds both the MC and AU power by the positive term $\\frac{\\Delta_\\theta\\Delta_Q}{\\Delta_\\theta+\\Delta_Q}(\\alpha^*_{\\mathrm{MS}}-\\alpha_{\\mathrm{MS}})$.","core_discovery":"The central claim is that preliminary misspecification testing—checking model assumptions before a model-based test—can be beneficial, contrary to much of the existing literature. The authors formalize a two-regime setup: data come either from a distribution $P_\\theta$ satisfying the model assumptions (where the constrained MC test has higher power) or from $Q$ violating them (where the unconstrained AU test has higher power), with $\\lambda$ the probability of the first regime. Under assumptions (I)–(IV), Lemma 1 establishes that there is a $\\lambda^*\\in(0,1)$ such that the combined procedure has strictly higher power than both unconditional tests at that mixture. The advantage over the better of the two unconditional tests equals $\\frac{\\Delta_\\theta\\Delta_Q}{\\Delta_\\theta+\\Delta_Q}(\\alpha^*_{\\mathrm{MS}}-\\alpha_{\\mathrm{MS}})>0$, where $\\Delta_\\theta$ and $\\Delta_Q$ are the power gaps and $\\alpha_{\\mathrm{MS}}$, $\\alpha^*_{\\mathrm{MS}}$ are the misspecification test's rejection probabilities under $P_\\theta$ and $Q$. Section 7 broadens this into the claim that model checking is more useful than the literature suggests, provided conditions (a)–(d) hold.","pith_inferences":["Extension: A practical corollary the authors leave implicit is that the misspecification test's significance level could be tuned to maximize the positive advantage term, since $\\alpha^*_{\\mathrm{MS}}-\\alpha_{\\mathrm{MS}}$ enters linearly; higher misspecification-test levels may be better when the cost of failing to switch to the AU test is large.","Extension: The $\\lambda$-mixture idea can be carried further: with a prior over a continuous family of violations, the same linearity argument would give conditions for the combined procedure to beat 'always use AU' or 'always use MC' in average power.","Extension: Because assumption (IV) is known to fail in common designs such as crossover trials, a quantitative bound on the power loss as the dependence parameter $\\delta$ grows is needed; the paper promises but does not deliver this, so the practical message currently rests on a plausibility argument.","Extension: The framework suggests a diagnostic for practice: before adopting a check-then-test protocol, simulate the two regimes and the misspecification test's sensitivity to see whether the four conditions hold for the specific tests and plausible violation distributions."],"forward_implications":["A combined procedure can strictly outperform both of its constituent tests in power for some mixture of valid and violated model assumptions, not just match the better one.","The conditions specify what makes model checking worthwhile: an informative misspecification test, approximate independence between the misspecification test and the main tests, and clear power advantages of both main tests in their own regimes.","Model checking should be aimed at detecting violations that are problematic for the main test, not at verifying that assumptions hold exactly.","Evaluating combined procedures only at $\\lambda=0$ or $\\lambda=1$, as most of the literature does, is too pessimistic; the mixture perspective can be more favorable.","In the two-sample normality example, the combined procedure performs close to the better of the t-test and the Wilcoxon-Mann-Whitney test, and can even exceed both in power for a skew-normal distribution, though its type 1 error can be mildly anti-conservative."],"supporting_citations":[{"why":"First demonstration that preliminary tests bias subsequent inference; the benchmark for combined-procedure analysis.","marker":"Bancroft (1944)"},{"why":"Asymptotic result that pretest procedures can only beat unconditional testing at the price of larger size; the main negative result the paper contests.","marker":"Albers et al. (2000b)"},{"why":"Simulation study concluding that normality pre-testing is not worthwhile for comparing two samples; a key representative of the pessimistic literature.","marker":"Rochon et al. (2012)"},{"why":"Survey of t-test versus Wilcoxon-Mann-Whitney and the interpretation of decision rules; supplies the AU test perspective and warns against normality testing.","marker":"Fay and Proschan (2010)"},{"why":"Coins the misspecification paradox that conditioning on passing a model check changes the conditional distribution; frames the paper's core conceptual issue.","marker":"Hennig (2007)"},{"why":"Argues misspecification and main tests pose different questions; the independence idea that backs assumption (IV).","marker":"Spanos (2010)"}],"fun_headline_variants":["Checking assumptions first can boost test power","Pre-testing model assumptions beats blind testing","When model checking improves test performance","Conditional testing: check assumptions, then gain power"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the independence of the misspecification test from both main tests (assumption (IV)); the authors call it very restrictive and unrealistic in most situations, and their promised relaxation to approximate independence with a small $\\delta$ is stated without proof.","fun_headline_variants_meta":{"raw":{"variants":["Checking assumptions first can boost test power","Pre-testing model assumptions beats blind testing","When model checking improves test performance","Conditional testing: check assumptions, then gain power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1409,"prompt_tokens":971,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":384}},"tokens_in":587,"tokens_out":438,"duration_ms":4839,"temperature":1.0,"reasoning_tokens":384,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:50:12.495253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a two-sample setup where the misspecification test and the main test are strongly dependent, for example the Shapiro-Wilk normality test and Welch's t-test on the same small normal sample, and compute the four conditional-probability deviations from independence in the paper's $\\delta$ relaxation. Then check whether at $\\lambda^*=\\Delta_Q/(\\Delta_\\theta+\\Delta_Q)$ the combined procedure's power still exceeds both unconditional powers; if it does not, condition (IV) is essential and the promised relaxation fails in that case.","supporting_citations":[{"cited_title":"BMC Medical Research Methodology 12:81--91","cited_arxiv_id":null,"evidence_quote":"Simulation study concluding that normality pre-testing is not worthwhile for comparing two samples; a key representative of the pessimistic literature."},{"cited_title":"Philosophia Mathematica 15:166--192","cited_arxiv_id":null,"evidence_quote":"Coins the misspecification paradox that conditioning on passing a model check changes the conditional distribution; frames the paper's core conceptual issue."},{"cited_title":"Journal of Econometrics 158:204--220","cited_arxiv_id":null,"evidence_quote":"Argues misspecification and main tests pose different questions; the independence idea that backs assumption (IV)."}],"review_version":1}