REVIEW 2 major objections 1 minor 86 references
Calibration without labels in multiple testing
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Pseudo-labels from p-value spacings allow calibration assessment of false discovery rate estimates without observing true labels.
desk verdict The pseudo-label construction from p-value spacings is a neat idea for label-free calibration checks, but the dependence from ordering makes the exact lfdr target claim non-obvious and in need of a clean derivation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Pseudo-labels derived from the spacings of ordered p-values, which serve as regression targets for the local false discovery rate.
What would settle it
A simulation in which the data-generating process is fully known so that true local false discovery rates can be computed exactly, yet the pseudo-labels fail to regress to those rates in expectation.
Extended reading notes
Core claim
We study how such claims can be interpreted as approximately calibrated forecasts of the null hypothesis, yielding interpretable error probabilities even under model misspecification. Our approach draws conceptual inspiration from probabilistic forecasting but addresses a different challenge: unlike forecasting, where labels are eventually observed, in multiple testing the ground truth is never revealed, so calibration must be assessed stochastically and established indirectly. We address this challenge by constructing a set of pseudo-labels, derived from the spacings of ordered p-values, which have the local false discovery rate as their regression target. Our construction unlocks existing
Load-bearing premise
The pseudo-labels built from p-value spacings have the local false discovery rate as their regression target even when the model is misspecified.
Editorial extensions
If this is right
- Standard calibration diagnostics from forecasting become applicable to multiple testing procedures.
- Error probabilities attached to individual hypotheses remain meaningful even when the underlying model is misspecified.
- Large empirical surveys of published work can detect miscalibration in common error measures such as the q-value.
- Post-hoc recalibration adjustments can be performed on existing multiple-testing outputs.
Reading between the lines
- The same spacing-based construction could be applied to other error measures besides the q-value.
- Routine reporting in empirical fields might eventually include calibration diagnostics derived from these pseudo-labels.
- The method opens a route to compare calibration properties across different multiple-testing procedures on the same data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a framework for assessing calibration of local false discovery rate (lfdr) estimates and related error measures like q-values in multiple testing without access to ground-truth labels. It constructs pseudo-labels from spacings of ordered p-values, asserts that these have the lfdr as their regression target (even under misspecification), and uses this to enable stochastic calibration checks. A large-scale empirical survey of psychology and neuroscience literature is presented as evidence that q-values can be severely miscalibrated.
Significance. If the pseudo-label construction is valid, the work supplies a practical tool for post-hoc calibration assessment in settings where labels are unavailable, extending ideas from probabilistic forecasting to multiple testing. The empirical survey finding, if reproducible, would be notable for fields that rely on FDR-based error measures.
major comments (2)
- [Abstract / core construction] The central technical claim is that pseudo-labels derived from spacings of the ordered p-values have the local FDR as their regression target (i.e., E[pseudo-label_i | p_i] = lfdr(p_i)). The skeptic note correctly identifies that each spacing depends on the full vector of order statistics, so the conditional expectation given only p_i does not automatically factor; an explicit derivation establishing the equality (under the paper's assumptions or under misspecification) is required before the calibration-assessment machinery can be used.
- [Empirical survey] The empirical claim that q-values are 'severely miscalibrated' rests on the pseudo-label calibration assessment. Because the validity of that assessment hinges on the unresolved conditional-expectation step above, the survey results cannot yet be interpreted as evidence of miscalibration.
minor comments (1)
- [Abstract] The abstract states the conceptual approach and the empirical finding but supplies no equations or proof sketches; readers cannot evaluate the construction from the given information.
Simulated Author's Rebuttal
We thank the referee for their careful and constructive review. We address the two major comments point by point below. We agree that an explicit derivation of the key conditional-expectation property is needed for clarity and will supply it in revision.
read point-by-point responses
-
Referee: [Abstract / core construction] The central technical claim is that pseudo-labels derived from spacings of the ordered p-values have the local FDR as their regression target (i.e., E[pseudo-label_i | p_i] = lfdr(p_i)). The skeptic note correctly identifies that each spacing depends on the full vector of order statistics, so the conditional expectation given only p_i does not automatically factor; an explicit derivation establishing the equality (under the paper's assumptions or under misspecification) is required before the calibration-assessment machinery can be used.
Authors: We accept that the manuscript would be strengthened by an explicit derivation. The pseudo-label construction uses normalized spacings of the ordered p-values; under the two-group model the joint distribution of the order statistics yields E[spacing-based pseudo-label_i | p_i] = lfdr(p_i) by direct computation of the conditional density. The same equality holds under misspecification because the target is defined as the regression function of the pseudo-label on p_i. We will add a self-contained appendix deriving the result from first principles and will reference it in the main text. revision: yes
-
Referee: [Empirical survey] The empirical claim that q-values are 'severely miscalibrated' rests on the pseudo-label calibration assessment. Because the validity of that assessment hinges on the unresolved conditional-expectation step above, the survey results cannot yet be interpreted as evidence of miscalibration.
Authors: We agree that interpretability of the survey depends on the core claim. Once the requested derivation is supplied, the large-scale survey of published q-values can be read as evidence that many reported q-values deviate from the pseudo-label regression target. The survey itself requires no substantive change; we will simply add a forward reference to the new appendix when discussing the empirical findings. revision: partial
Circularity Check
No significant circularity; derivation relies on external mathematical property of spacings
full rationale
The paper constructs pseudo-labels from spacings of ordered p-values and asserts they have local FDR as regression target. This is presented as a derived property rather than a definitional equivalence or fitted input renamed as prediction. No self-citation chains, ansatz smuggling, or uniqueness theorems from prior author work are invoked in the provided text to justify the central claim. The construction is independent of the target result and does not reduce to it by construction; the equality is an external claim subject to verification rather than tautological. The derivation chain remains self-contained against external benchmarks.
Assumptions & free parameters
assumptions (1)
- domain assumption Spacings of ordered p-values yield pseudo-labels whose regression target is the local false discovery rate.
Cite this review
Pith. "Pith review of Calibration without labels in multiple testing." pith.science (2026). https://pith.science/paper/ZFYQAHOI
@misc{pith2026260619737,
author = {Pith},
title = {Pith review of: Calibration without labels in multiple testing},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZFYQAHOI}},
note = {Machine review of arXiv:2606.19737}
}
abstract
Large-scale hypothesis testing supports probability claims about individual hypotheses, as in empirical Bayes methods for estimating local false discovery rates. We study how such claims can be interpreted as approximately calibrated forecasts of the null hypothesis, yielding interpretable error probabilities even under model misspecification. Our approach draws conceptual inspiration from probabilistic forecasting but addresses a different challenge: unlike forecasting, where labels are eventually observed, in multiple testing the ground truth is never revealed, so calibration must be assessed stochastically and established indirectly. We address this challenge by constructing a set of pseudo-labels, derived from the spacings of ordered $p$-values, which have the local false discovery rate as their regression target. Our construction unlocks existing tools for assessing and performing post-hoc calibration in multiple testing. Notably, we find on a large-scale empirical survey of published psychology and neuroscience literature that the $q$-value, a popular error measure based on the false discovery rate, can be severely miscalibrated.
Figures
Reference graph
Works this paper leans on
-
[1]
Biometrika , pages=
A frequentist local false discovery rate , author=. Biometrika , pages=. 2025 , publisher=
2025
-
[2]
and Xiang, Daniel and Fithian, William , TITLE =
Soloff, Jake A. and Xiang, Daniel and Fithian, William , TITLE =. Ann. Statist. , FJOURNAL =. 2024 , NUMBER =
2024
-
[3]
, TITLE =
Gneiting, Tilmann and Balabdaoui, Fadoua and Raftery, Adrian E. , TITLE =. J. R. Stat. Soc. Ser. B Stat. Methodol. , FJOURNAL =. 2007 , NUMBER =
2007
-
[4]
Compendium of Meteorology: Prepared under the Direction of the Committee on the Compendium of Meteorology , pages=
Verification of weather forecasts , author=. Compendium of Meteorology: Prepared under the Direction of the Committee on the Compendium of Meteorology , pages=. 1951 , publisher=
1951
-
[5]
and Tusher, Virginia , TITLE =
Efron, Bradley and Tibshirani, Robert and Storey, John D. and Tusher, Virginia , TITLE =. J. Amer. Statist. Assoc. , FJOURNAL =. 2001 , NUMBER =
2001
-
[6]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Estimation of a two-component mixture model with applications to multiple testing , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2016 , publisher=
2016
-
[7]
Biostatistics , volume=
False discovery rates: a new deal , author=. Biostatistics , volume=. 2017 , publisher=
2017
-
[8]
The Annals of Applied Statistics , pages=
An empirical Bayes mixture method for effect size and false discovery rate estimation , author=. The Annals of Applied Statistics , pages=. 2010 , publisher=
2010
Show all 86 references
-
[9]
A direct approach to false discovery rates , author=. J. Roy. Statist. Soc. Ser. B , volume=. 2002 , publisher=
2002
-
[10]
arXiv preprint arXiv:2402.08792 , year=
Interpretation of local false discovery rates under the zero assumption , author=. arXiv preprint arXiv:2402.08792 , year=
-
[11]
PLoS biology , volume=
Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature , author=. PLoS biology , volume=. 2017 , publisher=
2017
-
[12]
International conference on machine learning , pages=
Distribution-free calibration guarantees for histogram binning without sample splitting , author=. International conference on machine learning , pages=. 2021 , organization=
2021
-
[13]
Advances in Neural Information Processing Systems , volume=
Distribution-free binary classification: prediction sets, confidence intervals and calibration , author=. Advances in Neural Information Processing Systems , volume=
-
[14]
Proceedings of the 22nd international conference on Machine learning , pages=
Predicting good probabilities with supervised learning , author=. Proceedings of the 22nd international conference on Machine learning , pages=
-
[15]
Journal of Applied Meteorology and Climatology , volume=
On subjective probability forecasting , author=. Journal of Applied Meteorology and Climatology , volume=
-
[16]
From Probability to Statistics and Back: High-Dimensional Models and Processes--A Festschrift in Honor of Jon A
Smooth and non-smooth estimates of a monotone hazard , author=. From Probability to Statistics and Back: High-Dimensional Models and Processes--A Festschrift in Honor of Jon A. Wellner , volume=. 2013 , publisher=
2013
-
[17]
and Guntuboyina, Adityanand and Pitman, Jim , TITLE =
Soloff, Jake A. and Guntuboyina, Adityanand and Pitman, Jim , TITLE =. Electron. J. Stat. , FJOURNAL =. 2019 , NUMBER =
2019
-
[18]
Advances in multivariate statistical methods , pages=
Inference in exponential family regression models under certain shape constraints using inversion based techniques , author=. Advances in multivariate statistical methods , pages=. 2009 , publisher=
2009
-
[19]
Controlling the false discovery rate:
Benjamini, Yoav and Hochberg, Yosef , fjournal=. Controlling the false discovery rate:. J. Roy. Statist. Soc. Ser. B , volume=. 1995 , publisher=
1995
-
[20]
A simple sequentially rejective multiple test procedure , author=. Scand. Actuar. J. , pages=. 1979 , publisher=
1979
-
[21]
Genovese, Christopher and Wasserman, Larry , TITLE =. Ann. Statist. , FJOURNAL =. 2004 , NUMBER =
2004
-
[22]
Genovese, Christopher and Wasserman, Larry , TITLE =. J. R. Stat. Soc. Ser. B Stat. Methodol. , FJOURNAL =. 2002 , NUMBER =
2002
-
[23]
Journal of Statistical Planning and Inference , volume=
Resampling-based false discovery rate controlling multiple test procedures for correlated test statistics , author=. Journal of Statistical Planning and Inference , volume=. 1999 , publisher=
1999
-
[24]
Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach , author=. J. Roy. Statist. Soc. Ser. B , volume=
-
[25]
Bioinformatics , volume=
Identifying differentially expressed genes using false discovery rate controlling procedures , author=. Bioinformatics , volume=. 2003 , publisher=
2003
-
[26]
2025 , howpublished =
Ignatiadis, Nikolaos and Sen, Bodhisattva , title =. 2025 , howpublished =
2025
-
[27]
International conference on machine learning , pages=
On calibration of modern neural networks , author=. International conference on machine learning , pages=. 2017 , organization=
2017
-
[28]
Proceedings of Thirty Eighth Conference on Learning Theory , pages =
Can a calibration metric be both testable and actionable? , author =. Proceedings of Thirty Eighth Conference on Learning Theory , pages =. 2025 , volume =
2025
-
[29]
Proceedings of the 55th Annual ACM Symposium on Theory of Computing , pages=
A unifying theory of distance from calibration , author=. Proceedings of the 55th Annual ACM Symposium on Theory of Computing , pages=
-
[30]
The comparison and evaluation of forecasters , author=. J. Roy. Statist. Soc. Ser. D , volume=. 1983 , publisher=
1983
-
[31]
American Journal of Roentgenology , volume=
Radiologists' performance for differentiating benign from malignant lung nodules on high-resolution CT using computer-estimated likelihood of malignancy , author=. American Journal of Roentgenology , volume=. 2004 , publisher=
2004
-
[32]
, author=
Forecasting Precipitation in Percentages of Probability. , author=. Monthly Weather Review , volume=
-
[33]
Dawid, A. P. , TITLE =. J. Amer. Statist. Assoc. , FJOURNAL =. 1982 , NUMBER =
1982
-
[34]
Advances in neural information processing systems , volume=
Verified uncertainty calibration , author=. Advances in neural information processing systems , volume=
-
[35]
International Conference on Artificial Intelligence and Statistics , pages=
Mitigating bias in calibration error estimation , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2022 , organization=
2022
-
[36]
, author=
Measuring calibration in deep learning. , author=. CVPR workshops , volume=
-
[37]
American Journal of Medical Genetics Part B: Neuropsychiatric Genetics , volume=
Calibration of credibility of agnostic genome-wide associations , author=. American Journal of Medical Genetics Part B: Neuropsychiatric Genetics , volume=. 2008 , publisher=
2008
-
[38]
Jarosław Błasiok and Preetum Nakkiran , booktitle=. Smooth
-
[39]
, booktitle=
Okoroafor, Princewill and Kleinberg, Robert and Kim, Michael P. , booktitle=. Near-Optimal Algorithms for Omniprediction , year=
-
[40]
Lee, Donghwan and Huang, Xinmeng and Hassani, Hamed and Dobriban, Edgar , TITLE =. J. Mach. Learn. Res. , FJOURNAL =. 2023 , PAGES =
2023
-
[41]
Arrieta-Ibarra, Imanol and Gujral, Paman and Tannen, Jonathan and Tygert, Mark and Xu, Cherie , TITLE =. J. Mach. Learn. Res. , FJOURNAL =. 2022 , PAGES =
2022
-
[42]
Proceedings of Thirty Eighth Conference on Learning Theory , pages =
Truthfulness of Decision-Theoretic Calibration Measures , author =. Proceedings of Thirty Eighth Conference on Learning Theory , pages =. 2025 , editor =
2025
-
[43]
Proceedings of the Eighteenth International Conference on Machine Learning , pages =
Zadrozny, Bianca and Elkan, Charles , title =. Proceedings of the Eighteenth International Conference on Machine Learning , pages =. 2001 , isbn =
2001
-
[44]
, TITLE =
Pyke, R. , TITLE =. J. Roy. Statist. Soc. Ser. B , FJOURNAL =. 1965 , PAGES =
1965
-
[45]
B. Smooth. arXiv preprint arXiv:2309.12236 , year=
-
[46]
Applied statistics , pages=
Plotting p against x , author=. Applied statistics , pages=. 1983 , publisher=
1983
-
[47]
Monthly weather review , volume=
Some remarks on the reliability of categorical probability forecasts , author=. Monthly weather review , volume=
-
[48]
and Ghosh, Joydeep , TITLE =
Banerjee, Arindam and Merugu, Srujana and Dhillon, Inderjit S. and Ghosh, Joydeep , TITLE =. J. Mach. Learn. Res. , FJOURNAL =. 2005 , PAGES =
2005
-
[49]
Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining , pages=
Transforming classifier scores into accurate multiclass probability estimates , author=. Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining , pages=
-
[50]
Robbins, Herbert , TITLE =. Ann. Math. Statist. , FJOURNAL =. 1964 , PAGES =
1964
-
[51]
2010 , PAGES =
Efron, Bradley , TITLE =. 2010 , PAGES =
2010
-
[52]
BMC bioinformatics , volume=
A unified approach to false discovery rate estimation , author=. BMC bioinformatics , volume=. 2008 , publisher=
2008
-
[53]
Klaus, Bernd and Strimmer, Korbinian , TITLE =. J. SFdS , FJOURNAL =. 2011 , NUMBER =
2011
-
[54]
Rice, Kenneth and Spiegelhalter, David , TITLE =. Statist. Sci. , FJOURNAL =. 2008 , NUMBER =
2008
-
[55]
What should the genome-wide significance threshold be?
Panagiotou, Orestis A and Ioannidis, John PA , fjournal=. What should the genome-wide significance threshold be?. Int. J. Epidemiol. , volume=. 2012 , publisher=
2012
-
[56]
Barlow, R. E. and Bartholomew, D. J. and Bremner, J. M. and Brunk, H. D. , TITLE =. 1972 , PAGES =
1972
-
[57]
Grotzinger, S. J. and Witzgall, C. , TITLE =. Appl. Math. Optim. , FJOURNAL =. 1984 , NUMBER =
1984
-
[58]
Proceedings of the
Robbins, Herbert , TITLE =. Proceedings of the
-
[59]
Robbins, Herbert , TITLE =. Rev. Inst. Internat. Statist. , FJOURNAL =. 1963 , PAGES =
1963
-
[60]
Efron, Bradley , TITLE =. Ann. Statist. , FJOURNAL =. 2007 , NUMBER =. doi:10.1214/009053606000001460 , URL =
2007 doi
-
[61]
Efron, Bradley , TITLE =. Statist. Sci. , FJOURNAL =. 2008 , NUMBER =
2008
-
[62]
Empirical
Efron, Bradley and Tibshirani, Robert , fjournal=. Empirical. Genet. Epidemiol. , volume=. 2002 , publisher=
2002
-
[63]
A new vector partition of the probability score , author=. J. Appl. Meteorol. Climatol. , volume=
-
[64]
, TITLE =
Gneiting, Tilmann and Raftery, Adrian E. , TITLE =. J. Amer. Statist. Assoc. , FJOURNAL =. 2007 , NUMBER =
2007
-
[65]
Advances in large margin classifiers , volume=
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods , author=. Advances in large margin classifiers , volume=. 1999 , publisher=
1999
-
[66]
and Vohra, Rakesh V
Foster, Dean P. and Vohra, Rakesh V. , TITLE =. Biometrika , FJOURNAL =. 1998 , NUMBER =
1998
-
[67]
A proof of calibration via
Foster, Dean P , fjournal=. A proof of calibration via. Games Econ. Behav. , volume=. 1999 , publisher=
1999
-
[68]
Econometrica , volume=
A simple adaptive procedure leading to correlated equilibrium , author=. Econometrica , volume=. 2000 , publisher=
2000
-
[69]
Multicalibration: Calibration for the (
Hebert-Johnson, Ursula and Kim, Michael and Reingold, Omer and Rothblum, Guy , booktitle =. Multicalibration: Calibration for the (. 2018 , volume =
2018
-
[70]
Proceedings of the 8th Conference on Innovations in Theoretical Computer Science (ITCS) , year =
Inherent Trade-Offs in the Fair Determination of Risk Scores , author =. Proceedings of the 8th Conference on Innovations in Theoretical Computer Science (ITCS) , year =
-
[71]
Obtaining well calibrated probabilities using
Naeini, Mahdi Pakdaman and Cooper, Gregory and Hauskrecht, Milos , booktitle=. Obtaining well calibrated probabilities using
-
[72]
, title =
Samworth, Richard J. , title =. Proceedings of the International Congress of Mathematicians 2026 , year =. 2509.26040 , archivePrefix =
2026
-
[73]
On the theory of mortality measurement:
Grenander, Ulf , fjournal=. On the theory of mortality measurement:. Scand. Actuar. J. , volume=. 1956 , publisher=
1956
-
[74]
2014 , PAGES =
Groeneboom, Piet and Jongbloed, Geurt , TITLE =. 2014 , PAGES =
2014
-
[75]
Robertson, Tim and Wright, F. T. and Dykstra, R. L. , TITLE =
-
[76]
Nature Medicine , pages=
An atlas of exposome--phenome associations in health and disease risk , author=. Nature Medicine , pages=. 2026 , publisher=
2026
-
[77]
Jama , volume=
Evolution of reporting P values in the biomedical literature, 1990-2015 , author=. Jama , volume=
1990
-
[78]
The tight constant in the
Massart, Pascal , fjournal=. The tight constant in the. Ann. Probab. , pages=. 1990 , publisher=
1990
-
[79]
Proceedings of the American Mathematical Society , volume=
On an extension of the concept conditional expectation , author=. Proceedings of the American Mathematical Society , volume=
-
[80]
Conditional expectation given a -lattice and applications , author=. Ann. Math. Statist. , volume=
-
[81]
Bernoulli , volume=
Isotonic conditional laws , author=. Bernoulli , volume=
-
[82]
Construction and comparison of statistical models , author=. J. Roy. Statist. Soc. Ser. B , volume=
-
[83]
Monthly weather review , volume=
Verification of forecasts expressed in terms of probability , author=. Monthly weather review , volume=
-
[84]
Tony , TITLE =
Sun, Wenguang and Cai, T. Tony , TITLE =. J. Amer. Statist. Assoc. , FJOURNAL =. 2007 , NUMBER =
2007
-
[85]
Elicitation of personal probabilities and expectations , author=. J. Amer. Statist. Assoc. , volume=
-
[86]
A general method for comparing probability assessors , author=. Ann. Statist. , volume=
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.