{"id":"761a8795-2e02-43a3-9694-39397f3df73c","arxiv_id":"2607.15015","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Augmenting an omnibus test with conditionally calibrated secondary statistics and a small Type I error budget preserves primary power and sharply increases sensitivity to feature-specific departures.","lead":"This paper proposes a way to add extra checks to an existing goodness-of-fit test: keep the main test's rejection budget almost intact, but reserve a small slice for simple diagnostics like variance or skewness, with each extra check calibrated under the null after the earlier checks pass. The method is a single fixed rectangle rule once calibrated; a focused simulation shows that giving just 0.75% of the 5% budget to variance and skewness raises KS power against a narrow-sca","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the augmentation principle is internally sound and the experiment supports it; prespecification is explicitly required.","rationale":"The reader's ACCEPT verdict is well-supported. My independent read confirms that the central construction is logically sound: the budget identities (11)-(15) are correct, the order-invariance proposition is clearly qualified, and the Monte Carlo calibration consistency proof is standard. The experiment is carefully designed with paired standard errors, a full sensitivity grid, and explicit acknowledgement that the displayed allocation is illustrative, not optimal. The only caveat I see is the finite-sample accuracy of quantile estimates in the extreme tails of the secondary statistics, where the effective calibration sample is about 94 per tail. However, the paper quantifies this uncertainty (Appendix B) and empirically demonstrates that null size deviations are tiny across all allocations. The prespecification requirement is a standard and explicitly stated premise; it limits the applicability of the method but does not undermine the proof-of-principle claim. Therefore, I do not regard it as a load-bearing concern that should change the verdict.","tokens_in":35621,"tokens_out":15309,"duration_ms":166455,"concrete_test":"Re-run the main experiment with a smaller calibration sample (e.g., m=10^5 instead of 10^6) and check whether the null rejection probability remains within a few Monte Carlo standard errors of 0.05 and whether the power gain against N(0,0.2^2) stays large. If the size distorts materially, the extreme-tail calibration is less stable than reported; if not, the result is robust to finite-m effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After reviewing the construction, theory, and experiment, I find no load-bearing flaw in the central claim. The chain test is a fixed rectangular rule; conditional calibration correctly allocates the unconditional budgets so that the null rejection probability is exactly alpha at the population level (Section 2.4), and Proposition 3.1 gives a rigorous asymptotic justification for the Monte Carlo calibration. The finite-sample concern about calibrating at extremely small secondary levels (about 94 expected observations per tail) is addressed empirically: the reported null sizes are within 1.18 standard errors of 0.05 across the full allocation grid, and the power gains are far larger than calibration-to-calibration variation. The paper's claim is explicitly conditional on prespecification (Section 2.8), which is a standard requirement for any hypothesis test; the paper warns against post-hoc tuning and provides a sensitivity grid. No internal inconsistency or circular step was found. The reader's weakest assumption (prespecification) is a valid caveat but not a flaw in the argument, and it does not change the verdict.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an augmentation principle for goodness-of-fit testing: retain a primary omnibus statistic but allocate a small, prespecified fraction of the total null rejection budget to secondary statistics that capture simple feature-specific departures. The implementation is an ordered chain T0→...→TL in which each stage's acceptance region is calibrated under the null conditional on acceptance by all earlier stages. Unconditional stage budgets b_r are converted into conditional levels γ_r via Eq. (15), giving the exact population size identity α=1−∏(1−γ_r). The final test is a fixed rectangular acceptance rule; ordering is a boundary-selection and attribution device. Theoretical support includes strong consistency of the reusable Monte Carlo calibration (Prop. 3.1), an exact randomized finite-m pooled-rank variant (Prop. 3.2), a sufficient condition for order invariance (Prop. 2.1), and an algebraic equivalence to weighted minimum-p/hyperrectangle rules (Prop. 2.2). The experiment augments KS with variance and skewness under a standard normal null with n=10, m=h=10^6, reporting that transferring 0.75% of α=0.05 changes power against N(0.5,1) from 0.2742 to 0.2733 while raising power against N(0,0.2^2) from 0.4693 to 0.9827. A sensitivity grid over ρ and ω, null-rejection checks, and stagewise decompositions are provided.","tokens_in":35854,"tokens_out":19931,"duration_ms":218466,"significance":"If the result holds, the paper supplies a simple, transferable tool: rather than replacing an omnibus test, one can add a few interpretable diagnostics with a tightly controlled error-budget trade-off. The theoretical core is clean: the budget-to-conditional-level mapping (Eqs. (11)–(15)) is algebraically sound, and the proofs of Props. 2.1, 2.2, 3.1, and 3.2 are credible and appropriately conditional on standard regularity assumptions. Strengths include complete, reproducible R code, a prespecified sensitivity grid rather than post-hoc tuning, explicit treatment of finite-m calibration uncertainty, and plain acknowledgment of the proof-of-principle scope. The prespecification requirement (Section 2.8) is a standard and explicit condition, not an internal inconsistency. The reported size checks and paired comparisons make the power claims credible. The paper is not overclaimed: it repeatedly states that the result is not a uniform dominance proof and that the allocation illustrated is not an optimal budget. The methodology should be of interest to practitioners of goodness-of-fit testing and to statistical methodologists working on multi-statistic tests under a single null hypothes","major_comments":[],"minor_comments":[{"comment":"The phrase 'exact randomized finite-m null size' is correct, but the randomized character should be emphasized more prominently at first use. A reader might otherwise infer exactness in the ordinary nonrandomized sense; Section 3.3 itself is clear, but the conclusion restates only 'exact finite-m decision' without the qualifier.","section":"3.3 / abstract"},{"comment":"The weighted minimum-p baseline uses self-ranks r/m for calibration and (rank+1)/(m+1) for evaluation. This is asymptotically negligible, and the null-size check in Table 6 directly covers it, but one sentence explaining the finite-sample convention and why it does not affect the comparisons would improve transparency.","section":"4.2 / Appendix E"},{"comment":"The line 'Q(bRσ_m △ bRτ_m) → 0' is an abuse of notation: the probability Q is applied to the symmetric difference. Please write Q(bRσ_m △ bRτ_m)→0 or clarify that it is the Q-measure of the symmetric difference.","section":"2.6, after Prop. 2.1"},{"comment":"The caption 'Bold indicates the largest unrounded estimated power in a column; exact maxima are all shown in bold' is confusing because both clauses appear to say the same thing. It would be clearer to state simply that bold marks the maximum in each column, with exact unrounded values retained.","section":"Table 1 caption"},{"comment":"The paper repeatedly and correctly warns that ρ and ω must be prespecified. A short paragraph in the conclusion on possible principled, data-free selection heuristics—e.g., utility functions or worst-case power criteria over a prespecified alternative grid—would strengthen practical applicability. This is a scope suggestion, not a required fix.","section":"4.3.7 / 5"}],"recommendation":"minor_revision","confidential_remarks":"I found no load-bearing error. The prespecification concern raised by one reader is adequately addressed in Section 2.8 and should not block publication. The only minor issue I would flag for the editor is terminological: calling the construction 'priority-preserving' is accurate in terms of budget attribution, but since the final rule is a simultaneous rectangular intersection, the 'priority' is an interpretive property of the boundary-selection mechanism. This is already explained in the paper, but reviewers or readers familiar with gatekeeping methods might initially expect a different kind of priority. No action is required beyond what the authors already do."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the central mechanism is sound. The size identity alpha = 1 - prod(1 - gamma_r) and the conversion between unconditional budgets and conditional levels (eqs 11–15) are correct, and the power experiment is careful enough to believe the proof-of-principle claim. A tiny 0.75% transfer to variance and skewness buys large gains against scale and tail alternatives while barely denting KS power. That is a genuine, practically useful observation.\n\nWhat is actually new is the ordered conditional-calibration scheme: calibrate each secondary acceptance region under the null conditional on surviving earlier stages, and parameterize by unconditional first-rejection budgets. The resulting rule is a fixed hyperrectangle—Proposition 2.2 says so explicitly—but the boundary-selection logic and the ordered first-rejection decomposition are the contribution. The pooled-rank exact finite-m construction (Prop 3.2) is a nice extra, though it is for one observation at a time, and the paper knows its limits.\n\nThe experiment is a model of honest reporting: code shipped, paired standard errors, full sensitivity grid, null sizes within 1.18 SEs of nominal. I trust the numbers.\n\nSoft spots, in proportion. The load-bearing assumption is prespecification: the budgets rho and omega (and which secondary statistics, and their order) must be fixed before seeing data. The paper says this clearly in Section 2.8, and the sensitivity grid shows the trade-off, but it does not give a principled rule for choosing rho or omega. In practice people will tune them, and then the size guarantee silently breaks. That is a real limitation, though a standard one for hypothesis testing. Also, at the displayed allocation each secondary tail is calibrated on only about 94 expected observations; the paper acknowledges the finite-m effect but does not deeply study it beyond reporting SEs. And Proposition 2.1's order-invariance needs conditional independence, which likely does not hold for these statistics; the near-invariance in the tables is empirical, not proven.\n\nWho it's for: applied statisticians who like interpretable diagnostics bolted onto an omnibus test. It deserves a serious referee. I would send it out, expecting minor-to-moderate revisions on the prespecification discussion, not a rewrite.","headline":"Solid, honestly-scoped augmentation method: the calibration logic checks out and the experiment is careful, but prespecification of the budgets is doing real work and the new bottle is an old class of rules.","tokens_in":36316,"tokens_out":2956,"would_cite":true,"duration_ms":30405,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62F40"],"pacs":[],"model":"deepseek-v4-flash","headline":"Reserving a tiny slice of a goodness-of-fit test's Type I error budget for secondary diagnostics can sharply broaden its power without sacrificing the primary test's performance.","keywords":["goodness-of-fit testing","test augmentation","conditional calibration","Type I error allocation","power decomposition","Kolmogorov-Smirnov test","secondary statistics","exact finite-sample test"],"falsifier":"Simulate a practitioner who inspects a few samples from N(0,0.2^2) and then chooses rho and omega to maximize power, while still using the paper's calibration formulas on the same null bank. Under H0, measure the null rejection rate of this post-hoc-tuned chain: it should exceed the nominal 0.05, whereas the same chain with prespecified budgets stays at 0.05.","tokens_in":35494,"feed_emoji":"📊","tokens_out":3894,"duration_ms":42213,"temperature":0.7,"pith_summary":"This paper argues that an omnibus goodness-of-fit test can be kept as the primary decision rule and still gain sensitivity to specific departures, by reserving a small, prespecified fraction of its nominal Type I error probability for secondary diagnostics such as sample variance and skewness. The mechanism is an ordered chain in which each secondary acceptance region is calibrated under the null conditional on having passed all earlier stages, so the overall null rejection probability remains exactly alpha while the ordering supplies a clean accounting of what each added statistic contributes. In a normal-null experiment with n=10, handing 0.75% of the 0.05 budget to variance and skewness leaves power against a location shift essentially unchanged (0.2742 to 0.2733) yet lifts power against a scale alternative N(0,0.2^2) from 0.4693 to 0.9827. The paper supports the principle with strong consistency of the Monte Carlo calibration and an exact finite-m pooled-rank variant.","feed_headline":"Tiny error-budget shift gives goodness-of-fit tests big power gains","feed_subtitle":"A 0.75% slice of the null budget multiplies scale-detection power while leaving location power intact.","key_machinery":"The ordered conditional-calibration chain T0 -> T1 -> ... -> TL. Each acceptance region Br is chosen so that Q0(Tr not in Br | X in Ar-1) = gamma_r, with unconditional stage budgets b_r = (product_{s<r}(1-gamma_s)) gamma_r summing to alpha = 1 - product(1-gamma_r). This turns the design problem into a transparent choice of how much null rejection probability to move from primary to secondary diagnostics, and it turns the final test into a fixed rectangular acceptance region. Supporting machinery includes an order-invariance proposition under conditional null independence, a proof that weighted minimum-p tests are equivalent to hyperrectangle tests, a strong-consistency theorem for the Monte","core_discovery":"The central claim is the augmentation principle: you can broaden the power of an established test by spending a tiny, unconditional slice of its rejection budget on feature-specific statistics, with each secondary statistic's acceptance boundaries calibrated conditional on acceptance by all earlier stages. The resulting test is a fixed rectangular acceptance rule; the ordering is only a mechanism for choosing boundaries and assigning first-rejection credit, not a sequential-sampling scheme. The paper's unconditional budget parameterization b_0 + ... + b_L = alpha makes the trade-off explicit, and the experiment shows that transferring 0.75% of the total 0.05 budget from Kolmogorov-Smirnov to","pith_inferences":["The method suggests a general recipe for any omnibus test: preselect a small set of interpretable diagnostics and fix a tiny budget fraction before seeing data; the same conditional-calibration chain should work for multivariate or two-sample settings with permutation calibration.","The near order-invariance seen at small budgets implies that the choice of ordering mainly affects how credit is reported, not the decision; this could let practitioners choose order for interpretability without worrying about power changes.","A principled, data-independent rule for selecting rho and omega is still missing; the sensitivity grid is descriptive rather than an optimization, so the next step would be a utility- or worst-case-based budget selector.","The exact pooled-rank version may be especially useful when calibration reproducibility matters per observation, at the cost of recomputing boundaries for each new sample."],"forward_implications":["Any prespecified diagnostic — variance, skewness, tail indices, or problem-specific summaries — can be appended to any primary omnibus test without asymptotically inflating the null level beyond alpha.","The cost of augmentation is directly measurable in units of primary rejection probability (alpha - b0), and the benefit is measured stagewise by ordered first-rejection contributions under alternatives.","The weak equivalence of weighted minimum-p and hyperrectangle tests means the augmentation results transfer to simultaneous decision rules; the chain differs only in how coordinate thresholds are selected.","Under conditional null independence given primary acceptance, the order of secondary statistics does not change the final test, so attribution, not total power, is the order-sensitive output.","The pooled-rank variant gives exact finite-m null size for a fixed null law, with the trade-off that the boundaries become observation-specific rather than reusable."],"fun_headline_variants":["0.75% budget shift lifts scale power from 47 to 98","Tiny budget slice, huge tail-power boost in tests","Conditional calibration: 0.75% cost, 2x power","Spend 0.75% on diagnostics, keep KS power nearly intact"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"All secondary statistics, their tail rules, their order, and the budget fractions rho and omega must be fixed before looking at the data; if any of these are chosen after inspecting the sample, the claimed null rejection level is no longer guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["0.75% budget shift lifts scale power from 47 to 98","Tiny budget slice, huge tail-power boost in tests","Conditional calibration: 0.75% cost, 2x power","Spend 0.75% on diagnostics, keep KS power nearly intact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1254,"prompt_tokens":852,"completion_tokens":402,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":323}},"tokens_in":596,"tokens_out":402,"duration_ms":4767,"temperature":1.0,"reasoning_tokens":323,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:25:02.767668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a practitioner who inspects a few samples from N(0,0.2^2) and then chooses rho and omega to maximize power, while still using the paper's calibration formulas on the same null bank. Under H0, measure the null rejection rate of this post-hoc-tuned chain: it should exceed the nominal 0.05, whereas the same chain with prespecified budgets stays at 0.05.","supporting_citations":[],"review_version":1}