Pith. sign in

REVIEW 3 major objections 2 minor 2 cited by

Covariance-adaptive residual testing with stagewise critical values controls dependent multiple tests and recovers signals more cleanly than marginal methods.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 14:49 UTC pith:KOHV7A5G

load-bearing objection Abstract-only multiple-testing claim; the supplied full text is a different gr-qc paper, so we cannot audit the residual identities, admissibility argument, or simulations. the 3 major comments →

arxiv 2606.07466 v2 pith:KOHV7A5G submitted 2026-06-05 stat.ME

Covariance-Adaptive Residualization and Stagewise Calibration for Dependent Multiple Testing

classification stat.ME MSC 62H1562J15
keywords multiple testingcovariance-adaptive residualizationstagewise calibrationMRDprecision matrixFDRGaussian meansdependent tests
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When many means are tested at once and the observations share an arbitrary covariance structure, ordinary marginal p-values waste the dependence information. This paper keeps the Maximum Residual Down residualization idea of Cohen et al. but replaces its model-specific thresholds with the generalized step-down constants of Gavrilov et al., producing a simple, covariance-adaptive step-down rule. The resulting procedure is monotone residual-based, so earlier admissibility theory applies at once. The same residuals can be rewritten through a single active precision matrix, which both cuts computation and shows that residualization is geometry of the precision matrix restricted to the still-active coordinates. Simulations across many dependence patterns show lower normalized misclassification risk than popular marginal competitors; under several structured covariances the method simultaneously keeps FDR near the target, drives FNR nearly to zero, pushes power near one, and rejects roughly the true number of signals. Dependence is therefore not merely a nuisance: residualization plus stagewise calibration can turn it into a tool for cleaner large-scale inference.

Core claim

A stagewise-calibrated version of Maximum Residual Down residualization yields an admissible, covariance-adaptive step-down procedure for multivariate Gaussian means under arbitrary dependence; the residuals admit a single active-precision-matrix representation, and the resulting tests frequently dominate marginal methods on normalized misclassification risk while recovering nearly all signals under structured dependence.

What carries the argument

Covariance-adaptive residual statistics of the MRD type, re-expressed through one active precision matrix and calibrated by Gavrilov et al. generalized step-down critical constants; the construction places the rule inside the monotone residual-based step-down class whose admissibility is already known.

Load-bearing premise

The data are jointly Gaussian and the covariance (or its precision matrix) is known well enough that the residual statistics and their stagewise thresholds remain valid.

What would settle it

Under a non-Gaussian multivariate law, or with badly misspecified covariance, check whether the procedure still keeps FDR near the nominal level while matching or beating marginal competitors on normalized misclassification risk and signal recovery; systematic failure would overturn the performance claims.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The abstract claims a multiple-testing procedure for multivariate Gaussian means under arbitrary covariance: it keeps the covariance-adaptive residualization of Cohen et al.’s Maximum Residual Down (MRD) method, replaces MRD’s model-dependent thresholds by stagewise critical constants from Gavrilov et al., asserts admissibility by membership in the monotone residual-based step-down class of Ghosh and Chakrabarti (2026), derives an active-precision-matrix representation of the residuals that reduces computation, and reports simulations in which the method often has lower normalized misclassification risk than marginal procedures and, under structured dependence, simultaneously near-nominal FDR, very small FNR, power near one, and rejections near the true signal count. The supplied full-text block, however, is not that manuscript: it is the complete gr-qc paper arXiv:2606.07467 (“Stochastic scalar-tensor inflation and beyond”), with no theorems, residual definitions, precision-matrix identities, admissibility arguments, or simulation tables for the multiple-testing claims.

Significance. If the abstract’s claims were substantiated in a matching manuscript, the work would be of clear interest in large-scale dependent testing: a covariance-adaptive residual step-down with a simple calibration rule, an admissibility link to a general monotone class, a computational reduction via active precision geometry, and strong finite-sample recovery under structured dependence would be useful contributions. Those strengths cannot be credited or audited from the material provided, because the body of the submission is a different paper in a different field.

major comments (3)
  1. Title/abstract vs. full text: the abstract and paper_id identify a stat.ME multiple-testing paper (MRD residualization, Gavrilov stagewise constants, Ghosh–Chakrabarti admissibility, active precision matrix, FDR/FNR/power simulations). The full manuscript text is instead Launay’s gr-qc paper on stochastic scalar-tensor inflation (EFToDE, Mukhanov–Sasaki sources, Horndeski/EGB examples). No equation, section, table, or proof from the claimed paper is present. The central claims therefore cannot be checked for correctness, and the submission as supplied is not reviewable as the stated work.
  2. Admissibility (abstract): the abstract states that admissibility “follows directly” from membership in the monotone residual-based step-down class of Ghosh and Chakrabarti (2026). That reference shares year and (per the reader’s note) authorship with the present work. Without the actual definitions, class membership proof, and error-rate control arguments, it is impossible to verify that the construction is not circular or that the cited general theory applies under the paper’s residual and calibration choices.
  3. Gaussianity and known/usable covariance (abstract): performance and residualization are stated for multivariate Gaussian means under arbitrary covariance, with residuals tied to an active precision matrix. The formal assumptions, estimation of dependence, and robustness when covariance is misspecified or non-Gaussian are not available in the supplied text, so the load-bearing conditions for FDR control and the reported simulation superiority cannot be assessed.
minor comments (2)
  1. Once the correct full manuscript is supplied, the abstract’s simulation claims (normalized misclassification risk, simultaneous FDR/FNR/power/rejection counts under structured dependence) will need to be tied to explicit tables, dependence models, and competitors; those materials are absent here.
  2. The abstract’s “alternative representations … through a single active precision matrix” should be matched to numbered identities and complexity statements in the correct paper; they cannot be located in the gr-qc text provided.

Circularity Check

1 steps flagged

Admissibility is imported wholesale from a same-author 2026 citation; residualization and stagewise calibration themselves are not shown to be circular from the available text.

specific steps
  1. self citation load bearing [Abstract (admissibility claim)]
    "Since the resulting procedure belongs to the class of monotone residual-based step-down procedures studied by Ghosh and Chakrabarti (2026), its admissibility follows directly from their general theory."

    The paper’s only stated theoretical guarantee of admissibility is not proved here; it is imported from a same-author, same-year citation. If that prior work’s uniqueness/admissibility theorem is itself built around residual step-down constructions of the same type, the present claim is load-bearing on an unverified self-citation chain rather than an independent derivation. The abstract does not exhibit the membership proof, so the reduction cannot be checked equation-by-equation, but the logical structure is: define procedure → assert membership in authors’ own class → conclude admissibility.

full rationale

The only text belonging to arXiv:2606.07466 is the abstract (the supplied full manuscript is a mismatched gr-qc paper on stochastic inflation). From that abstract, the central theoretical claim—admissibility of the proposed procedure—is not derived in-paper; it is asserted to follow immediately once the procedure is placed in the monotone residual-based step-down class of Ghosh and Chakrabarti (2026), i.e., the same authors. That is a load-bearing self-citation for uniqueness/admissibility. Nothing in the abstract exhibits a definitional loop for the residual statistics, the active-precision-matrix representation, the Gavrilov-style stagewise constants, or the simulation comparisons to marginal procedures; those pieces cite external work (Cohen et al. 2009; Gavrilov et al. 2009) or are empirical. Without the correct full text one cannot check whether membership in the monotone residual class is automatic by construction or a nontrivial verification. Proportionate score: self-citation carries the admissibility claim (pattern 3/4), but the methodological and simulation content is not shown to reduce to its own inputs. Score 4, not higher.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

Abstract-only ledger. The procedure rests on multivariate Gaussian means, a usable covariance/precision structure, residualization as in MRD, and step-down constants as in Gavrilov et al., plus membership in the monotone residual step-down class of Ghosh and Chakrabarti (2026). No free parameters or invented physical entities are named in the abstract; simulation designs and any estimated covariance are unknown here.

axioms (4)
  • domain assumption Observations are multivariate Gaussian with means under simultaneous test and arbitrary but exploitable covariance dependence.
    Stated scope of the abstract; residualization and precision-matrix geometry rely on this model.
  • domain assumption The proposed rule is a monotone residual-based step-down procedure in the sense of Ghosh and Chakrabarti (2026), hence admissible.
    Admissibility is not re-proved; it is imported from the authors’ general theory paper.
  • domain assumption Gavrilov et al. (2009) generalized step-down critical constants provide a valid stagewise calibration for the residual statistics under the stated dependence.
    Calibration rule is taken from that literature; abstract does not re-derive error control.
  • domain assumption MRD residualization (Cohen et al. 2009) remains valid when thresholds are replaced by the stagewise rule and residuals are rewritten via a single active precision matrix.
    Core mechanism retained from prior work; computational rewrite is claimed equivalent.

pith-pipeline@v1.1.0-grok45 · 30411 in / 2728 out tokens · 28313 ms · 2026-07-12T14:49:26.855900+00:00 · methodology

0 comments
read the original abstract

In this paper, we study simultaneous hypothesis testing for multivariate Gaussian means under arbitrary covariance dependence. Building upon the Maximum Residual Down (MRD) procedure of Cohen et al. (2009), we investigate a systematic stagewise calibration strategy based on the generalized step-down critical constants of Gavrilov et al. (2009). The proposed methodology retains the covariance-adaptive residualization mechanism of MRD while replacing the original model-dependent threshold specification with a simple and principled calibration rule. Since the resulting procedure belongs to the class of monotone residual-based step-down procedures studied by Ghosh and Chakrabarti (2026), its admissibility follows directly from their general theory. We also derive alternative representations of the MRD residual statistics that express all active residuals through a single active precision matrix, substantially reducing computational complexity while revealing a direct connection between covariance-adaptive residualization and active precision-matrix geometry. Extensive simulation studies under a broad range of dependence structures demonstrate that the proposed methodology frequently achieves substantially lower normalized misclassification risk than several widely used marginal testing procedures. Under several structured dependence models, it also exhibits remarkably strong signal-recovery behavior, simultaneously attaining false discovery rates close to the nominal level, extremely small false non-discovery rates, powers approaching one, and average numbers of rejections close to the expected number of true signals. These findings suggest that covariance-adaptive residualization and stagewise calibration play complementary roles in exploiting dependence information for large-scale multiple testing under arbitrary covariance structures.

Figures

Figures reproduced from arXiv: 2606.07466 by Arijit Chakrabarti, Prasenjit Ghosh.

Figure 1
Figure 1. Figure 1: Normalized misclassification rates (NMR) for the competing multiple [PITH_FULL_IMAGE:figures/full_fig_p022_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Empirical false discovery rates (FDR) for the competing multiple testing [PITH_FULL_IMAGE:figures/full_fig_p026_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Empirical false non-discovery rates (FNR) for the competing multiple [PITH_FULL_IMAGE:figures/full_fig_p028_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Empirical power of the competing multiple testing procedures under six [PITH_FULL_IMAGE:figures/full_fig_p029_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Average numbers of rejections (ANR) produced by the competing multi [PITH_FULL_IMAGE:figures/full_fig_p031_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Normalized misclassification rates (NMR) for the competing multiple [PITH_FULL_IMAGE:figures/full_fig_p055_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Empirical false discovery rates (FDR) for the competing multiple test [PITH_FULL_IMAGE:figures/full_fig_p059_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Empirical false non-discovery rates (FNR) for the competing multiple [PITH_FULL_IMAGE:figures/full_fig_p064_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Empirical power of the competing multiple testing procedures under six [PITH_FULL_IMAGE:figures/full_fig_p068_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Average numbers of rejections (ANR) produced by the competing [PITH_FULL_IMAGE:figures/full_fig_p072_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Asymptotic Bayes Optimality Under Sparsity of the Gavrilov-Benjamini-Sarkar Step-Down Testing Procedure

    math.ST 2026-07 accept novelty 6.5

    The GBS step-down FDR procedure is asymptotically Bayes-optimal under sparsity over a broad class of sparse Gaussian regimes.

  2. Bayesian Model Pursuit and Near-Oracle Sparse Signal Discovery Under Dependence

    stat.ME 2026-06 unverdicted novelty 6.0

    BSD achieves near-oracle Bayes risk and support recovery for sparse signals under arbitrary known covariance dependence via posterior-guided model pursuit.

Reference graph

Works this paper leans on

9 extracted references · cited by 2 Pith papers

  1. [1]

    or could generate increased complexity beyond the single-field scenario, as a field space curvature might be generated in the multifield case [111]. In the aforementioned existing calculations, beyond the decoupling limit or not, several approximations are employed on top of perturbativity and the tree-level limitation �, such as the freezing of couplings...

  2. [2]

    We do not make any claims on the validity of those approximations

    3) the linear regime ofζon those scales (equating to using eqn.(2.20) in the limit of 1) and 2)). We do not make any claims on the validity of those approximations. Given the un- certainty on the matter from the cited work, we bring to the attention of the reader that, indeed some non-Gaussianity would arise in the tail of the distribution even within thi...

  3. [3]

    Identify linear gauge-invariants such as QI �δϕ I + ˙ϕI b H Φ (4.5a) and pick one to coarse-grain (hereQ I, following a generalisation of our approach with� but that is not compulsory).Q I has a known matrix-mass Mukhanov-Sasaki equation of the form � � t QI + 3H� tQI � � � a� QI +� I J QJ = 0,(4.6) where� t(�)I �∂ t(�)I + ΓI JK ˙ϕJ b (�)K is the field-sp...

  4. [4]

    Express all perturbations of one gauge in terms of one of them (QI). Using the con- straints with their new stress-energy contributions δρ= G IJ ˙ϕI b δ ˙ϕJ +∂ I V δϕ I �G IJ ˙ϕI b ˙ϕJ b Ψ,(4.7a) δ� i =�G IJ ˙ϕI b ∂iδϕJ .(4.7b) These terms are particularly special because they, apart from the∂ I Vcontribution, constitute the adiabatic contributions (meani...

  5. [5]

    Obtain the sources to the nonlinear equations , eqn. (4.4), by linearising them and ap- plyingQ I � �Q J> �W I J QJ in the previously found decompositions, whereWis a win- dowing matrix allowing to coarse-grain depending on the crossing of each field. Given our previous findings, sources should all be found proportional to that of eqn. (4.6). The final eq...

  6. [6]

    Starobinsky,� ��� ���� �� ��������� ������������ ������ ������� �����������,������� ������� ���(1980) 99

    A. Starobinsky,� ��� ���� �� ��������� ������������ ������ ������� �����������,������� ������� ���(1980) 99

  7. [7]

    Guth,����������� ��������� � �������� �������� �� ��� ������� ��� ������� ��������, �������� ������ ���(1981) 347

    A.H. Guth,����������� ��������� � �������� �������� �� ��� ������� ��� ������� ��������, �������� ������ ���(1981) 347

  8. [8]

    Linde,� ��� ����������� �������� ��������� � �������� �������� �� ��� �������� �������� ������������ �������� ��� ���������� �������� ��������,������� ������� ����(1982) 389

    A. Linde,� ��� ����������� �������� ��������� � �������� �������� �� ��� �������� �������� ������������ �������� ��� ���������� �������� ��������,������� ������� ����(1982) 389

  9. [9]

    Mukhanov and G.V

    V.F. Mukhanov and G.V. Chibisov,������� ������������ ��� � ����������� ��������, ���� �������(1981) 532. – 30 –