REVIEW 3 major objections 5 minor 37 references
A nonparametric statistical method for deconvolving densities in the analysis of proteomic data
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper introduces NPFD, a nonparametric deconvolution method that raises the ratio of estimated Fourier transforms to the Nth power to estimate a target density fY, and reports that this stable, smooth estimator handles large variance…
desk verdict Genuinely new deconvolution construction with strong simulation results, but the theoretical target is a scaled sum of N copies, not fY, and the paper overstates the case. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the N-Power Fourier Deconvolution (NPFD) transform: for a chosen integer N, the samples from the convolving and mixed distributions are linearly transformed with a = 1/√N and shifts derived from their means, so that the quotient of the estimated Fourier transforms equals the Fourier transform of a scaled-shifted copy of Y. The Nth power of this quotient is then computed, which exponentially suppresses the oscillating tails that plague ordinary Fourier deconvolution. The power N is selected automatically by increasing N until the transformed Fourier estimate falls below a threshold ε, and the final density estimate is obtained by Fourier inversion with Monte Carlo integration.
What would settle it
Simulate many datasets from a known target density that is not closed under convolution, such as a Gumbel distribution, with a normal convolving density, and apply NPFD for N = 1, 2, 4, 8, 16 while holding everything else fixed. If the integrated squared error grows steadily with N even though the N = 1 estimate is accurate, then the exponentiation is estimating a different object and the claim that NPFD estimates fY fails for large N.
Extended reading notes
Core claim
The central discovery is that deconvolution can be stabilized by raising the quotient of Fourier transforms to a power N. For independent variables with fZ = fX * fY, the identity (φ_Z̃(t)/φ_X̃(t))^N = φ_{aΣY_k + (1−√N)E[Y]}(t) holds, where Y_1,...,Y_N are i.i.d. copies of Y, a = 1/√N, and the shift is chosen so that the scaled sum has the same mean and variance as Y. Thus, taking the Nth power of the estimated ratio gives the characteristic function of a scaled and shifted sum of N copies of the target variable, not of Y itself except in convolution-closed families. The paper argues that inverting this Nth-powered ratio suppresses high-frequency noise in the Fourier tails and yields a smooth, accurate estimate of fY across a wide range of distributions, including the Gumbel example that is not closed under convolution.
Load-bearing premise
For N greater than 1, NPFD estimates the density of a scaled sum of N independent copies of the target variable, and the paper assumes that this sum density is close enough to the target density itself for a wide range of distributions.
Editorial extensions
If this is right
- NPFD would provide a stable deconvolution option in genomics and proteomics, where the variance of the observed mixture often substantially exceeds that of the target, a setting where classical methods such as FDD tend to oscillate.
- The same framework would apply both to measurement-error models and to separating two convolved signal densities, removing the need for separate algorithms for these two scenarios.
- By avoiding kernel bandwidth selection, NPFD would replace one tuning choice with an automatic selection of the power N and threshold ε.
- On the GerontoSys proteomic data, NPFD produces smooth estimates of the pure extrinsic aging signal and finds a statistically significant difference between age groups, whereas FDD produces unusable oscillating estimates on the same data.
- If the method is adopted, density deconvolution results in applications with high noise variance could become reproducible without strong parametric assumptions about the error distribution.
Reading between the lines
- A testable consequence that the paper leaves implicit is that the accuracy of NPFD should degrade as N grows for target distributions that are far from convolution-closed, because the Nth power estimates the density of a scaled sum of N copies of Y rather than Y itself; for such targets, an error-versus-N curve would reveal whether the smoothing is buying bias instead of stability.
- The repeated-measurement version of NPFD inherits the symmetry and centering assumptions on the error distribution from the estimator used to generate error samples, so its practical reliability in heteroscedastic settings is an open extension rather than an established property.
- The N-power damping of Fourier tails could in principle be combined with existing kernel deconvolution methods to estimate the density of the scaled N-sum directly and then solve for fY, but the paper does not develop that composite estimator.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces NPFD, a nonparametric deconvolution method that applies the N-th power of the ratio of estimated Fourier transforms of transformed samples from the mixture and convolving distributions. The claim is that, after appropriate scaling and shifting, this power reduces numerical instability and yields accurate estimates of the target density even when the convolving distribution has much larger variance than the target. The method is evaluated in extensive simulations against several existing deconvolution procedures (FDD, MCD, DKM, RMD, WEPF) and applied to Framingham blood-pressure data and GerontoSys proteomic skin-aging data. The paper includes an R package on CRAN.
Significance. If the central claim were fully established, NPFD would fill a practical gap: most existing nonparametric deconvolution methods degrade when the mixture variance dominates, and NPFD offers a simple Fourier-based alternative with smooth density estimates. As a computational proposal, the method is attractive and the experimental comparison is broad, covering multiple competing methods with 500 replications in each scenario. The real data applications, especially the proteomic extrinsic-aging deconvolution, are relevant and demonstrate useful behavior. However, the theoretical target of the estimator is not the density fY of the target variable but the density of a scaled and shifted N-fold convolution power of Y (Eq. 10). The paper does not provide a theorem, bias bound, or consistency result establishing that this convolution-power density is close to fY for general (non-convolution-closed) target distributions. This missing theoretical support is load-bearing because the method's name and abstract claim accurate estimation of fY, and the automatic power-selection rule does not control the shape distortion introduced by convolution.
major comments (3)
- [Section 4.1, Eq. (10)-(12)] The central derivation correctly shows that the N-th power of the Fourier ratio equals the characteristic function of S_N = N^{-1/2} \sum_{k=1}^N Y_k + (1-\sqrt{N})E(Y), not of Y itself. The inverse Fourier transform in Eq. (12) therefore yields f_{S_N}, not f_Y, for N>1. The paper states in Section 4.1 that equality holds only for convolution-closed families, and in Section 5.2.1 it is forced to set Nmax=1 for a two-mode normal mixture because convolving that density with itself changes the number of modes. Yet the abstract and Section 1 assert that NPFD leads to 'accurate estimation of fY'. No theorem, quantitative bound, or simulation isolating the approximation error f_{S_N} - f_Y is provided. This is the load-bearing gap: without a bound or a reinterpretation of the estimand, the claim that NPFD estimates fY is unsupported for general target distributions. Please add a formal statement of the approximation target and either a bias bound or a carefully designed simulation that directly measures the convolution-power bias for non-convolution-closed targets.
- [Section 4.2, Algorithm C1] The automatic selection of N is based solely on the decay of |\hat{\varphi}_{\tilde Y}(t)|^N below a threshold ε at a point t_k and t_k+δ. This criterion does not measure, or even bound, the distance between the density of S_N and the targeted f_Y. For a skewed or multimodal target, increasing N always reduces the characteristic-function tail and will be selected, while simultaneously increasing the shape bias induced by taking the N-fold convolution. Thus the algorithm can systematically select a power that oversmooths the target. The procedure would benefit from a diagnostic that monitors the change in the inverted density as N grows, or from an explicit bias-variance trade-off in the choice of N. At a minimum, the paper should acknowledge that the selection rule targets numerical stability, not fidelity to f_Y.
- [Section 5, simulation design] The simulation scenarios are predominantly smooth, unimodal, or nearly normal-like targets (Gamma, chi-square, Laplace convolutions, and a normal mixture only when N is artificially capped at 1). The Gumbel target in Scenarios 7-8 of Section 5.1.1 is a non-convolution-closed exception, and NPFD reports moderate ISE values (median 10×ISE 0.12-0.35), but no scenario isolates the structural bias f_{S_N} - f_Y from sampling variability and from the smoothing introduced by the Poisson-based density pre-estimation. As a result, the claimed wide applicability of the method to general densities is not empirically supported. Please add at least one simulation with a strongly skewed, bimodal, or otherwise non-convolution-closed target where N>1 is actually used and where the error is decomposed into the convolution-power bias and the smoothing/estimation error.
minor comments (5)
- [Section 5.1.1, first paragraph] The text says 'we used in in the estimation of fX three instead of five degrees of freedom' - there is a duplicated 'in'. Also, the sentence structure is awkward; please rephrase.
- [Section 5.1.2, last paragraph of Section 5.1] In the sentence 'Neumann [20] utilized, in contrast to Diggle and Hall [19], the exact representation of the Fourier transform of the mixed density to compute the smoothing kernel', the contrast is not fully clear; consider adding a clause explaining that FDD uses a damping function estimated from the data, whereas MCD uses the known analytic form of the Fourier transform in its kernel construction.
- [Section 4.2, Algorithm C1] The algorithm breaks out of the inner loop when |\hat{\varphi}_{\tilde Y}(t_k)|^N > 1, but it does not explicitly state that this case increments N and restarts; the accompanying text explains this, but the pseudocode could include a comment to make the control flow unmistakable.
- [Section 6.2] The caption of Figure 7 correctly states that the older age group has four women, but the main text says 'for the five women of the younger age group and four women of the older age group'; the number of older women is indeed four because one proband was excluded, but the sentence 'we removed all proteins that were not available in at least three of the samples in at least one of the four groups' could be misread as implying an equal sample size per group; please rephrase for clarity.
- [Appendix B, density estimation] The method of Wand [37] for choosing the number of histogram bins is said to 'make use of kernel density estimation', but the specific form of the MISE-minimizing criterion is not given; a one-sentence recap of the Wand rule would help readers who are not familiar with that reference.
Circularity Check
No significant circularity: NPFD's estimand is a scaled N-fold sum density, but that is a correctness gap, not a circular derivation.
full rationale
The derivation chain is self-contained and does not reduce to its inputs. Given the convolution model (3), φ_Y = φ_Z / φ_X is standard Fourier deconvolution (Eq. 4), and the N-power step in Eq. (10) explicitly states that (φ_{tilde Z}/φ_{tilde X})^N = φ_{a∑Y_k + (1−√N)E(Y)}. The inversion (12) is therefore exactly the density of S_N = N^{-1/2}∑Y_k + (1−√N)E(Y), not a hidden repackaging of f_Y. The power N is chosen by an internal threshold on the estimated Fourier transform (Section 4.2, Algorithm C1); no parameter of f_Y is fitted, and no quantity later called a prediction is an input. The cited uniqueness theorem is Rudin's external Fourier uniqueness theorem, not an author-imported theorem. The only self-citations are [28], a Poisson-regression density pre-smoothing modification used in implementation, and [33], the CRAN package; neither supplies the load-bearing claim that the N-th power recovers f_Y. The acknowledged limitation in Section 4.1, that the scaled-sum density equals f_Y only for convolution-closed families, and the Nmax = 1 workaround for the two-mode normal mixture in Section 5.2.1 are precisely a soundness/consistency gap: no bias bound is given for general f_Y. That is a correctness risk rather than circularity, because the estimand is openly derived from Eq. (10) and is not defined to be f_Y.
Assumptions & free parameters
free parameters (6)
- Power N =
data-dependent, e.g., 1 to 100
- Threshold epsilon =
0.001 default, 0.1 or nx^{-1/2} for small samples
- Margin delta =
small, not concretely specified
- Number of Fourier points K and upper bound t_K =
chosen suitably, no explicit rule
- Degrees of freedom for the natural cubic spline =
5 (default), 3 in Scenarios 1 and 2 of Section 5.1.1
- Number of histogram bins (Wand's method) =
adaptive to data
assumptions (4)
- standard math Fourier transform inversion theorem and convolution theorem
- domain assumption Additive model Z = X + Y with independent components
- ad hoc to paper The scaled-sum density f_S is a good approximation to fY
- domain assumption Var(Z) > Var(X) for a meaningful deconvolution
Cite this review
Pith. "Pith review of A nonparametric statistical method for deconvolving densities in the analysis of proteomic data." pith.science (2026). https://pith.science/paper/LBRMPFHH
@misc{pith2026250601540,
author = {Pith},
title = {Pith review of: A nonparametric statistical method for deconvolving densities in the analysis of proteomic data},
year = {2026},
howpublished = {\url{https://pith.science/paper/LBRMPFHH}},
note = {Machine review of arXiv:2506.01540}
}
abstract
In medical research, often, genomic or proteomic data are collected, with measurements frequently subject to uncertainties or errors, making it crucial to accurately separate the signals of the genes or proteins, respectively, from the noise. Such a signal separation is also of interest in skin aging research in which intrinsic aging driven by genetic factors and extrinsic, i.e.\ environmentally induced, aging are investigated by considering, e.g., the proteome of skin fibroblasts. Since extrinsic influences on skin aging can only be measured alongside intrinsic ones, it is essential to isolate the pure extrinsic signal from the combined intrinisic and extrinsic signal. In such situations, deconvolution methods can be employed to estimate the signal's density function from the data. However, existing nonparametric deconvolution approaches often fail when the variance of the mixed distribution is substantially greater than the variance of the target distribution, which is a common issue in genomic and proteomic data. We, therefore, propose a new nonparametric deconvolution method called N-Power Fourier Deconvolution (NPFD) that addresses this issue by employing the $N$-th power of the Fourier transform of transformed densities. This procedure utilizes the Fourier transform inversion theorem and exploits properties of Fourier transforms of density functions to mitigate numerical inaccuracies through exponentiation, leading to accurate and smooth density estimation. An extensive simulation study demonstrates that NPFD effectively handles the variance issues and performs comparably or better than existing deconvolution methods in most scenarios. Moreover, applications to real medical data, particularly to proteomic data from fibroblasts affected by intrinsic and extrinsic aging, show how NPFD can be employed to estimate the pure extrinsic density.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Deconvolution analysis of hormone data
V eldhuis, JD and Johnson, ML. Deconvolution analysis of hormone data. Methods in Enzymology, 210:539–575, 1992
work page 1992
-
[2]
Cdseqr: fas t complete deconvolution for gene expression data from bulk tissues
Kang, K, Huang, C, Li, Y, Umbach, DM, and Li, L. Cdseqr: fas t complete deconvolution for gene expression data from bulk tissues. BMC Bioinformatics, 22:262, 2021
work page 2021
-
[3]
Density deconvolu tion with laplace errors and unknown variance
Cai, J, Horrace, WC, and Parmeter, CF. Density deconvolu tion with laplace errors and unknown variance. Journal of Productivity Analysis, 56:103–113, 2021
work page 2021
-
[4]
L., Byers, J., Cabral, S., Celi, L
Charpignon, M. L., Byers, J., Cabral, S., Celi, L. A., Fer nandes, C., Gallifant, J., Lough, M. E., Mlombwa, D., Moukheiber, L., Ong, B. A., Panitchote, A., William, W ., Wong, A. I., and Nazer, L. Critical bias in critical care devices. Critical Care Clinics, 39(4):795–813, Oct 2023
work page 2023
-
[5]
Information bias in health research: Defi nition, pitfalls, and adjustment methods
Althubaiti, A. Information bias in health research: Defi nition, pitfalls, and adjustment methods. Journal of Multidisciplinary Healthcare, 9:211–217, May 2016
work page 2016
-
[6]
Badrick, T. Biological variation: Understanding why it is so important? Practical Laboratory Medicine , 23:e00199, Jan 2021
work page 2021
-
[7]
Banna, J. C., Fialkowski, M. K., and Townsend, M. S. Misre porting of dietary intake affects estimated nutrient in- takes in low-income spanish-speaking women. Journal of the Academy of Nutrition and Dietetics , 115(7):1124– 1133, Jul 2015
work page 2015
-
[8]
B., Eekhout, I., Boers, M., van der Vleuten, C
Mokkink, L. B., Eekhout, I., Boers, M., van der Vleuten, C . P . M., and de V et, H. C. W . Studies on reliability and measurement error of measurements in medicine – from design to statistics explained for medical researchers. Patient Related Outcome Measures, 14:193–212, Jul 2023
work page 2023
Show all 37 references
-
[9]
Wan, E. Y . F., Y u, E. Y . T., Chin, W . Y ., Fong, D. Y . T., Choi,E. P . H., and Lam, C. L. K. Association of blood pressure and risk of cardiovascular and chronic kidney disease in hong kong hypertensive patients. Hypertension, 74(2):331–340, Aug 2019
2019
-
[10]
Systematic er rors in peptide and protein identification and quantifi- cation by modified peptides
Bogdanow, B., Zauber, H., and Selbach, M. Systematic er rors in peptide and protein identification and quantifi- cation by modified peptides. Molecular & Cellular Proteomics , 15(8):2791–2801, Aug 2016
2016
-
[11]
R. J. Carroll, D. Ruppert, L. A. Stefanski, and C. Craini ceanu. Measurement Error in Nonlinear Models: A Modern Perspective. Chapman & Hall/CRC, 2 edition, 2006
2006
-
[12]
Density estimation with het eroscedastic error
Delaigle, A and Meister, A. Density estimation with het eroscedastic error. Bernoulli, 14(2):562–579, 2008. 23 Anarat et al
2008
-
[13]
On deconvolution w ith repeated measurements
Delaigle, A, Hall, P, and Meister, A. On deconvolution w ith repeated measurements. The Annals of Statistics , 36(2):665–685, 2008
2008
-
[14]
Density estimation in the p resence of heteroscedastic measurement error of un- known type using phase function deconvolution
Nghiem, L and Potgieter, CJ. Density estimation in the p resence of heteroscedastic measurement error of un- known type using phase function deconvolution. Statistics in Medicine , 37:3679–3692, 2018
2018
-
[15]
Deconvolution estimation in measu rement error models: The r package decon
Wang, XF and Wang, B. Deconvolution estimation in measu rement error models: The r package decon. Journal of Statistical Software, 39(10):1–24, 2011
2011
-
[16]
Delaigle and P
A. Delaigle and P . Hall. Methodology for non-parametri c deconvolution when the error distribution is unknown. Journal of the Royal Statistical Society: Series B (Statist ical Methodology), 78(1):231–252, 2016
2016
-
[17]
Comparative analysis of cell mixtures deconvolut ion and gene signatures generated for blood, immune and cancer cells
Alonso-Moreda, N, Berral-González, A, De La Rosa, E, Go nzález-V elasco, O, Sánchez-Santos, JM, and De Las Rivas, J. Comparative analysis of cell mixtures deconvolut ion and gene signatures generated for blood, immune and cancer cells. International Journal of Molecular Scienc...
2023
-
[18]
Proteome-wide analysis reveals an age-associated cellular phenotype of in situ age d human fibroblasts
Waldera-Lupa, DM, Kalfalah, F, Florea, AM, Sass, S, Kru se, F, and Rieder, V et al. Proteome-wide analysis reveals an age-associated cellular phenotype of in situ age d human fibroblasts. Aging (Albany NY) , 6(10):856– 878, 2014
2014
-
[19]
A fourier approach to nonparamet ric deconvolution of a density estimate
Diggle, PJ and Hall, P. A fourier approach to nonparamet ric deconvolution of a density estimate. Journal of the Royal Statistical Society: Series B (Methodological) , 55(2):523–531, 1993
1993
-
[20]
On the effect of estimating the error densi ty in nonparametric deconvolution
Neumann, MH. On the effect of estimating the error densi ty in nonparametric deconvolution. Journal of Non- parametric Statistics, 7(4):307–330, 1997
1997
-
[21]
Real and Complex Analysis
Rudin, W . Real and Complex Analysis . McGraw-Hill, New Y ork, 3rd edition, 1987
1987
-
[22]
F ourier Transforms
Bochner, S and Chandrasekharan, K. F ourier Transforms. Princeton University Press, Princeton, 1950
1950
-
[23]
On the optimal rates of convergence for nonparam etric deconvolution problems
Fan, J. On the optimal rates of convergence for nonparam etric deconvolution problems. The Annals of Statistics , 19(3):1257–1272, 1991
1991
-
[24]
Numerical calculation of stable densities a nd distribution functions
Nolan, JP. Numerical calculation of stable densities a nd distribution functions. Communications in Statistics - Stochastic Models, 13(4):759–774, 1997
1997
-
[25]
and Shanthikumar, J
Shaked, M. and Shanthikumar, J. G. Stochastic Orders . Springer Science+Business Media, New Y ork, 1st edition, 2007
2007
-
[26]
Bootstrap bandwidth select ion in kernel density estimation from a contaminated sample
Delaigle, A and Gijbels, I. Bootstrap bandwidth select ion in kernel density estimation from a contaminated sample. Annals of the Institute of Statistical Mathematics , 56(1):19–47, 2004
2004
-
[27]
Using specially designed ex ponential families for density estimation
Efron, B and Tibshirani, R. Using specially designed ex ponential families for density estimation. The Annals of Statistics, 24:2431–2461, 1996
1996
-
[28]
Empirical bayes analysis of single nucleotide polymorphisms
Schwender, H and Ickstadt, K. Empirical bayes analysis of single nucleotide polymorphisms. BMC Bioinfor- matics, 9(1), 2008
2008
-
[29]
Springer Berlin Heidelberg, Berlin, Heidelberg, 2015
H Wo´ zniakowski.Monte Carlo Integration. Springer Berlin Heidelberg, Berlin, Heidelberg, 2015
2015
-
[30]
deconvolve: Deconvol ution tools for measurement error problems
A Delaigle, T Hyndman, and T Wang. deconvolve: Deconvol ution tools for measurement error problems. https://github.com/TimothyHyndman/deconvolve, 2024. R package version 0.1.0
2024
-
[31]
The roles of ise and mise in density estimatio n
Jones, MC. The roles of ise and mise in density estimatio n. Statistics & Probability Letters, 12(1):51–56, 1991
1991
-
[32]
Functional k-sample problem when data are density functions
Delicado, P. Functional k-sample problem when data are density functions. Computational Statistics, 22:391– 410, 2007
2007
-
[33]
NPFD: N-Power F ourier Deconvolution, 2024
Akin Anarat. NPFD: N-Power F ourier Deconvolution, 2024. R package version 1.0.0
2024
-
[34]
Deconvoluting kernel den sity estimators
Stefanski, L and Carroll, RJ. Deconvoluting kernel den sity estimators. Statistics, 21:169–184, 1990
1990
-
[35]
Periodogram analysis and continuous spe ctra
Bartlett, MS. Periodogram analysis and continuous spe ctra. Biometrika, 37(1-2):1–16, 1950
1950
-
[36]
On optimal and data-based histograms
Scott, DW. On optimal and data-based histograms. Biometrika, 66:605–610, 1979
1979
-
[37]
Data-based choice of histogram bin width
Wand, MP. Data-based choice of histogram bin width. The American Statistician , 51:59–64, 1997. 24 Anarat et al. Appendix This appendix provides additional material that supports t he main manuscript. Appendix A provides a detailed de- scription of the deconvolution methods ou...
1997
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.