Pith. sign in

REVIEW 5 major objections 5 minor 29 references

Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees

T0 review · 5 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that a single Polya-tree shrinkage prior, applied through a hierarchical Bayes model, yields more accurate estimates of high-dimensional generalized linear model coefficients than standard penalized and empirical-Bayes…

desk verdict A promising Polya-tree shrinkage method for large-scale inference, but the flagship logistic regression claim rests on single-realization simulations and needs replication and sensitivity analysis before it can be taken as stated. read the letter →

arxiv 1908.08444 v5 pith:XR6SVSRC submitted 2019-08-22 stat.ME

classification stat.ME MSC 62J0762C1262G0562F15
keywords PolyatreehierarchicalBetamodelshrinkageestimationgeneralizedlinearmodelsempiricalBayescompounddecisiontheorynonparametricdeconvolutionlogisticregression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

In high-dimensional regression with fixed (nonrandom) coefficients, the paper asks for one shrinkage estimator that automatically adapts to the coefficient distribution—whether sparse, Gaussian, or mixed. It first shows that if the empirical distribution of the true coefficients were known, the Bayes rule against that distribution would be the optimal separable rule for the whole vector, in both frequentist and Bayesian senses. To approach this oracle without knowing the coefficients, the paper models the coefficients as i.i.d. draws from a common distribution and assigns that distribution a Polya-tree (hierarchical Beta) prior. The posterior mean of this mixing distribution estimates the empirical CDF of the coefficients nonparametrically, and the posterior means of the coefficients serve as shrinkage estimates. In three simulated high-dimensional logistic regressions, this method attains lower mean squared error for coefficients, linear predictor, and success probabilities than the MLE, the debiased MLE, LASSO, and Ridge.

What carries the argument

The central object is the L-level hierarchical Beta prior, which is a Polya tree on the space of coefficient distributions: it generates a piecewise-constant density on a fixed grid by recursively splitting intervals, where each split multiplies a parent probability by an independent Beta(1,1) variable. This construction gives conjugacy in the no-noise case, so posterior conditional probabilities are simple Beta updates, and it permits a Gibbs sampler for noisy and general likelihoods. In the logistic regression setting, each coefficient is updated conditionally on the others with a density proportional to the likelihood times the step-function prior, evaluated on a dense grid. The machinery does two jobs at once: the posterior mean of the mixing distribution performs nonparametric deconvolution, estimating the empirical CDF of the true coefficients, and the posterior means of the coefficients provide the shrinkage estimates that mimic the oracle rule.

What would settle it

Repeat Example 1 with the true nonzero coefficients moved to ±30, outside the [-24, 24] grid, while keeping all other settings fixed; if the hierarchical Beta estimates fail to shrink toward the true support or their MSE rises above LASSO's, the claimed universality would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that one nonparametric hierarchical Bayes construction—the L-level hierarchical Beta model, a Polya tree with Beta(1,1) splits—can serve as a universal shrinkage prior for high-dimensional regression. The theoretical anchor is an oracle result: under a symmetric sequence model with nonrandom parameters, the Bayes estimator against the empirical distribution of the parameters is instance-optimal among separable decision rules, and its dependence on the unknowns is only through that empirical distribution. The hierarchical Beta prior is proposed as a fully Bayes way to estimate the mixing distribution and hence approximate the oracle. In the logistic regression application, the posterior mean of the coefficient vector, computed by a Gibbs sampler that conditions each coefficient on the others, shrinks the MLE toward the estimated mixing distribution, and in the three reported simulation examples it achieves the lowest mean squared error among the five methods compared, for the coefficients, for the linear predictor, and for the success probabilities.

Load-bearing premise

The method's practical success depends on the user choosing a fixed grid wide and fine enough to contain the true coefficients—the simulations do this by setting endpoints such as -24 and 24—and the paper provides no data-driven rule or sensitivity analysis for this choice, while the proved optimality applies to the oracle rule that already knows the coefficient distribution rather than to the actual posterior-mean estimator.

Editorial extensions

If this is right

  • Across the three simulated logistic regressions, the hierarchical Beta posterior means achieve the lowest normalized MSE among the MLE, adjusted MLE, LASSO, Ridge, and the proposed method, for the coefficient vector, the linear predictor, and the success probabilities.
  • The same Gibbs-sampler recipe, with no per-example tuning beyond the grid and level count, yields marginal 95% credible intervals with observed coverage between 92.5% and 99% in the three examples.
  • In the compound-decision theory section, the optimal separable rule for fixed nonrandom coefficients depends on the unknowns only through their empirical CDF, which is exactly what the posterior mean of the Polya-tree mixing distribution estimates.
  • In the accident-data application, the hierarchical Beta posterior mean for Poisson rates has lower risk than the NPMLE plug-in estimate and nearly matches two oracle estimators, while all four Bayesian estimators beat the MLE by roughly a factor of four.
  • Because the same prior and sampler are applied verbatim across very different logistic examples, the paper suggests the method is a single recipe rather than a tuned-per-problem procedure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's fixed-grid choice (e.g., endpoints -24 and 24 in the logistic examples) is untested for sensitivity; a natural extension is an adaptive or expanding grid that tracks the support of the estimated mixing distribution, which would make the method genuinely tuning-free.
  • The oracle equivalence suggests a benchmark for other nonparametric Bayesian priors on the mixing distribution, such as Dirichlet-process mixtures or log-spline priors, since any prior that consistently estimates the empirical CDF should approach the same oracle.
  • Because the method estimates the full mixing distribution, it could be lifted to tasks the paper does not pursue, such as variable selection, ranking, or treatment-effect heterogeneity, by using posterior functions of the coefficients rather than just their means.
  • The reported MSE gains in Section 6.2 are based on single simulation realizations; averaging over repeated draws of the design and outcomes would give standard errors and clarify how much of the apparent margin is sampling noise, particularly in the example where the gap over Ridge is small.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a hierarchical Bayes approach to large-scale inference in which the unknown coefficient vector is modeled as i.i.d. draws from a common mixing distribution, and the mixing distribution is assigned a hierarchical Beta (Polya-tree-type) prior supported on a fixed grid. Section 2 gives a compound-decision justification for the oracle Bayes rule in the symmetric sequence model, showing that the oracle separable rule minimizes the compound risk for every fixed parameter vector. Section 3 defines the L-level hierarchical Beta model, derives the no-noise posterior, and provides Gibbs samplers for the sequence model and for logistic regression. Sections 4 and 5 apply the method to normal-means multiple testing and to the accident data of Simar et al. (1976). Section 6 applies the method to three high-dimensional logistic regression examples from Sur and Candès (2019), reporting in Table 3 that the hierarchical Beta method achieves the smallest normalized MSE for beta, mu, and q in each example. The abstract claims 'better estimation and prediction accuracy' compared with parametric and nonparametric alternatives.

Significance. If the empirical claims are taken at face value, the paper offers a practically useful nonparametric shrinkage method for high-dimensional GLMs, going beyond sequence-model deconvolution, and the Section 2 oracle optimality result is a clean contribution to compound decision theory. The accident-data analysis is a strength: it reports risk estimates with standard errors based on 400 replications. However, the central new claim about GLM performance rests on single-realization simulations without standard errors, and the theory in Section 2 does not cover the implemented GLM procedure. The paper also provides no code or sensitivity analysis for the key tuning inputs. Thus the broad claims in the abstract are not currently supported, although the underlying methodology is plausible and the issues appear addressable.

major comments (5)
  1. [Section 6.2, Table 3] The load-bearing empirical claim that 'The hierarchical Beta method performs best in each of the examples and for each choice of the model parameters' is based entirely on a single realization of X and Y for each example. No standard errors, no repeated replications, and no code are provided, so the reader cannot assess whether the reported ordering is stable or is an artifact of one favorable dataset. The authors should report results over many independent replications (with standard errors or boxplots) for all three examples, and should temper the abstract's general claim unless the replicated results support it.
  2. [Section 2 versus Section 6] The optimality result in Section 2 is established only for separable decision rules in the symmetric sequence model, as made explicit in equations (7)-(9). The hierarchical Beta posterior mean for the logistic regression coefficients in Section 6 is not a separable rule: the posterior for each beta_i depends on the entire vector Y and the full design matrix X through the conditional distribution in equation (23). Consequently, the Section 2 theory does not provide theoretical support for the GLM implementation, and the paper should either extend the theory or explicitly state that the GLM method is heuristic, with the sequence-model theory serving only as motivation.
  3. [Abstract and Section 6.2] The abstract claims 'better estimation and prediction accuracy,' but Table 3 measures only in-sample estimation error: normalized MSE for beta, mu, and q on the same data used to fit the model. There is no held-out prediction experiment, no cross-validation, and no evaluation on independent test data. The term 'prediction' is therefore unsupported, and the claims should be restricted to estimation accuracy unless genuine out-of-sample experiments are added.
  4. [Section 6.2 and Section 7] The method depends on user-specified inputs: the grid endpoints amin = -24, amax = 24, number of levels L = 6, the dense grid of K = 1280 values used in Algorithm 2, and the MCMC chain lengths. No sensitivity analysis is given for any of these choices, and Section 7 states only that an R package and vignette 'will provide guidance' in the future. Without a sensitivity study or a data-driven rule, the claim that the method is 'free of tuning' (Section 3 intro and Section 7) is overstated, and the possibility remains that performance degrades substantially for other reasonable grid choices, especially if the true coefficients lie outside [amin, amax].
  5. [Algorithm 2] Algorithm 2 as printed is incomplete: the loop over i in lines 3-6 computes values of f(b_k|pi) and f(y|b_k, beta_(i)) but never specifies how beta_i^(g) is actually sampled from its conditional distribution. This makes the main simulation results impossible to reproduce from the manuscript text alone. The authors should provide a complete, unambiguous pseudocode and, ideally, working code.
minor comments (5)
  1. [Abstract] The abstract says the prior 'assigns equal mass to every permutation of the fixed coefficient vector,' which is a useful intuition, but the body of the paper formalizes this only in the symmetric sequence model; the phrasing should be aligned with the formal scope.
  2. [Introduction, page 2] There is a typo: 'accodring' should be 'according.'
  3. [Section 6.2] The text says Examples 1 and 2 used a single MCMC simulation of 1000 iterations while Example 3 used 10 simulations of 150 iterations. No convergence diagnostics are reported, and the discrepancy across examples is not explained. A brief justification or trace plots would help.
  4. [Table 3] Table 3 reports MSE values only as fractions of the MLE MSE. Absolute MSE values and the scale of the target quantities would make the comparisons more interpretable, especially since beta and q are on very different scales.
  5. [Section 5] The accident-data analysis is one of the more compelling parts of the paper because it includes risk estimates with standard errors from 400 replications. Consider presenting this type of replication detail in the logistic regression section as a model for the missing analysis.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the oracle benchmark and the hierarchical Beta estimator are distinct objects, and the empirical comparisons are not defined in terms of their own outputs.

full rationale

The paper's theoretical optimality claim is explicitly restricted to the oracle separable Bayes rule under the symmetric sequence model, while the proposed hierarchical Beta estimator is a distinct, data-driven procedure that is not fitted to that oracle benchmark. No equation in Section 2 is used to define the hBeta posterior mean, and no fitted parameter is renamed as a prediction. The empirical superiority claim in Section 6.2 is a direct comparison of MSEs on simulated examples; the hierarchical Beta method is not constructed to minimize those MSEs, and the grid and level choices are analyst-specified tuning inputs rather than parameters fitted to the target quantities. The paper's selective-inference argument borrows from Yekutieli (2012), a co-author's prior work, but that citation supplies an external published mathematical framework rather than a conclusion whose only support is the present paper; moreover, the authors indicate a direct proof is available. The admitted limitations—single-realization simulation numbers, no sensitivity analysis for the grid, no data-driven grid-selection rule, and an unfinished R package—are correctness and reproducibility concerns, not circularity. The derivation chain therefore does not reduce to its inputs, and the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 1 invented entities

The central performance claim depends on modeling choices that are not derived: the Beta(1,1) hierarchy, the fixed grid with chosen endpoints and L levels, the Gibbs run lengths, and the correctness of the logistic likelihood. None of these are fitted to data in the formal sense, but they are analyst-chosen inputs. The oracle optimality result is standard compound decision theory and is not circular.

free parameters (4)
  • Number of levels L = 6, 7, or 8 depending on example
    Controls the resolution of the piecewise-constant mixing density; chosen per application in Sections 4, 5, and 6.
  • Grid endpoints (amin, amax) = [-5,5], [0,4], [-24,24]
    Must cover the support of the true coefficient distribution; set by the analyst for each example, with no sensitivity analysis.
  • Gibbs iterations and burn-in = G=1000 with 100 burn-in; 10 runs of 150 in Example 3
    MCMC run lengths are chosen per example; convergence is not formally diagnosed.
  • Dense grid K for beta sampling in Algorithm 2 = 1280 = 64*200 points
    Discretizes the conditional posterior of each beta coefficient; the approximation error is not reported.
assumptions (4)
  • domain assumption The coefficient vector beta is exchangeable and modeled as i.i.d. from a common unknown distribution
    Definition 2 and the hierarchical model in Section 3.3 impose exchangeability despite the stated frequentist fixed-effects view; this is central to the Bayesian pooling.
  • domain assumption The logistic likelihood with design matrix X is correctly specified
    All simulation results assume the generative logistic model in Definition 2; no robustness to misspecification is checked.
  • ad hoc to paper Beta(1,1) priors on the Polya tree split probabilities are appropriate
    Chosen for conjugacy and uniformity; no data-driven or theoretical justification beyond convenience.
  • ad hoc to paper The true mixing distribution is well approximated by a step function on a fixed grid
    The support and resolution of the prior are fixed by L and (amin, amax); the paper shows simulations only where this approximation holds.
invented entities (1)
  • Hierarchical Beta (Polya tree) prior on the mixing distribution
    purpose: Provides a nonparametric random-effects model for the unknown coefficient distribution and shrinks posterior means toward a learned density.
    The prior is a new statistical construct; its usefulness is demonstrated only in simulations and one reanalysis, with no external falsifiable prediction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees." pith.science (2026). https://pith.science/paper/XR6SVSRC

@misc{pith2026190808444,
  author       = {Pith},
  title        = {Pith review of: Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XR6SVSRC}},
  note         = {Machine review of arXiv:1908.08444}
}
abstract

Regularization in fitting regression models has been a highly active topic of research in the past few decades, but most of the existing methods are designed for particular situations, e.g. for the case of a sparse coefficient vector. We consider the problem of designing $\textit{universally}$ optimal regularized estimators in a given generalized linear model with fixed effects. First, we propose as a contender the Bayes estimator against an $\textit{ideal}$ prior that assigns equal mass to every permutation of the fixed coefficient vector, thus depending on the true coefficients only through their empirical CDF. We prove some optimality properties of this oracle estimator in both the frequentist and Bayesian frameworks. To compete with the oracle estimator, we posit a hierarchical Bayes model where the individual coefficients are modeled as i.i.d. draws from a common distribution $\pi$, which is in turn assigned a Polya tree prior to reflect indefiniteness. We demonstrate in examples that the posterior mean of $\pi$ under the postulated model adapts nonparametrically to the empirical CDF of the true coefficients. Correspondingly, the posterior means of the coefficients themselves are used to mimic the ideal estimator. Numerical experiments show that our method has better estimation and prediction accuracy compared to various parametric and nonparametric alternatives, from relatively standard $L_p$-regularized estimators to modern penalized-likelihood and Bayesian estimators for high dimensional regression.

Figures

Figures reproduced from arXiv: 1908.08444 by the authors.

Figure 1
Figure 1. Estimates of the success probabilities in a logistic regression simulated example. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Schematic of the hierarchical Beta model, [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. (a) Estimates of the local false discovery rate, fdr( [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Point estimates and interval estimates for selected observations. The curves are [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Estimates of success probabilities qi . Black circles correspond to the MLE. Top row is for Example 2, bottom row is for Example 3. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]
Figure 6
Figure 6. Figure 6: Estimates of model coefficients βj . 27 [PITH_FULL_IMAGE:figures/full_fig_p027_6.png]
Figure 7
Figure 7. Figure 7: Standard deviation of density estimators for [PITH_FULL_IMAGE:figures/full_fig_p031_7.png]
Figure 8
Figure 8. Figure 8: 15-level Hierarchical Beta density estimates, [PITH_FULL_IMAGE:figures/full_fig_p032_8.png]
Figure 9
Figure 9. Figure 9: Distribution of the deconvolution interval probabilities for the truncated [PITH_FULL_IMAGE:figures/full_fig_p033_9.png]
Figure 10
Figure 10. Figure 10: Deconvolution density estimates for the truncated-Normal Uniform mixing [PITH_FULL_IMAGE:figures/full_fig_p034_10.png]
Figure 11
Figure 11. Figure 11: Deconvolution estimates for the CDF of the truncated-Normal Uniform mix [PITH_FULL_IMAGE:figures/full_fig_p035_11.png]
Figure 12
Figure 12. Figure 12: Estimates of the CDF of the the posterior distribution of Θ [PITH_FULL_IMAGE:figures/full_fig_p036_12.png]
Figure 13
Figure 13. Figure 13: Marginal 95% credible intervals for the model coefficients [PITH_FULL_IMAGE:figures/full_fig_p037_13.png]
Figure 14
Figure 14. Figure 14: Gibbs posterior samples for πL. Top, middle and bottom rows correspond to Examples 1, 2 and 3, respectively. Left column: boxplots of the distribution of the 64 interval probabilities in Gibbs sampler runs 101 to 1000: π (101) i , ..., π (1000) i , for i = 1, ..., 64.…
Figure 15
Figure 15. Figure 15: Estimated mixing distribution CDF for the accident data of [PITH_FULL_IMAGE:figures/full_fig_p039_15.png]
Figure 16
Figure 16. Figure 16: CDF of estimated mixing distribution for simulated data. the blue curves are ()() [PITH_FULL_IMAGE:figures/full_fig_p040_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 28 canonical work pages

  1. [1]

    Benjamini and D

    Y. Benjamini and D. Yekutieli. False discovery rate--adjusted multiple confidence intervals for selected parameters. Journal of the American Statistical Association, 100 0 (469): 0 71--81, 2005

  2. [2]

    L. D. Brown and E. Greenshtein. Nonparametric empirical bayes and compound decision approaches to estimation of a high-dimensional vector of normal means. The Annals of Statistics, pages 1685--1704, 2009

  3. [3]

    B. P. Carlin and T. A. Louis. Bayes and empirical Bayes methods for data analysis. Chapman and Hall/CRC, 2010

  4. [4]

    L. H. Dicker and S. D. Zhao. High-dimensional classification via nonparametric empirical bayes and maximum likelihood inference. Biometrika, 103 0 (1): 0 21--34, 2016

  5. [5]

    B. Efron. Microarrays, empirical bayes and the two-groups model. Statistical science, 23 0 (1): 0 1--22, 2008

  6. [6]

    B. Efron. Tweedies formula and selection bias. Journal of the American Statistical Association, 106 0 (496): 0 1602--1614, 2011

  7. [7]

    B. Efron. Large-scale inference: empirical Bayes methods for estimation, testing, and prediction, volume 1. Cambridge University Press, 2012

  8. [8]

    B. Efron. Empirical bayes deconvolution estimates. Biometrika, 103 0 (1): 0 1--20, 2016

Show all 29 references
  1. [9]

    Efron, R

    B. Efron, R. Tibshirani, J. D. Storey, and V. Tusher. Empirical bayes analysis of a microarray experiment. Journal of the American statistical association, 96 0 (456): 0 1151--1160, 2001

  2. [10]

    Friedman, T

    J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of statistical software, 33 0 (1): 0 1, 2010

  3. [11]

    Hewitt and L

    E. Hewitt and L. J. Savage. Symmetric measures on cartesian products. Transactions of the American Mathematical Society, 80 0 (2): 0 470--501, 1955

  4. [12]

    Jiang and C.-H

    W. Jiang and C.-H. Zhang. General maximum likelihood empirical bayes estimation of normal means. The Annals of Statistics, 37 0 (4): 0 1647--1684, 2009

  5. [13]

    Kiefer and J

    J. Kiefer and J. Wolfowitz. Consistency of the maximum likelihood estimator in the presence of infinitely many incidental parameters. The Annals of Mathematical Statistics, pages 887--906, 1956

  6. [14]

    Koenker and I

    R. Koenker and I. Mizera. Convex optimization, shape constraints, compound decisions, and empirical bayes rules. Journal of the American Statistical Association, 109 0 (506): 0 674--685, 2014

  7. [15]

    Kwon and Z

    Y. Kwon and Z. Zhao. Nonparametric empirical bayes simultaneous estimation for multiple variances. arXiv preprint arXiv:1806.06377, 2018

  8. [16]

    D. V. Lindley and A. F. Smith. Bayes estimates for the linear model. Journal of the Royal Statistical Society. Series B (Methodological), pages 1--41, 1972

  9. [17]

    S. Reid, J. Taylor, and R. Tibshirani. Post-selection point and interval estimation of signal sizes in gaussian samples. Canadian Journal of Statistics, 45 0 (2): 0 128--148, 2017

  10. [18]

    H. Robbins. Asymptotically subminimax solutions of compound statistical decision problems, 1951. URL http://projecteuclid.org/euclid.bsmsp/1200500224

  11. [19]

    H. Robbins. An empirical bayes approach to statistics. Herbert Robbins Selected Papers, pages 41--47, 1956

  12. [20]

    Simar et al

    L. Simar et al. Maximum likelihood estimation of a compound poisson process. The Annals of Statistics, 4 0 (6): 0 1200--1209, 1976

  13. [21]

    C. M. Stein. Confidence sets for the mean of a multivariate normal distribution. Journal of the Royal Statistical Society. Series B (Methodological), pages 265--296, 1962

  14. [22]

    J. D. Storey. A direct approach to false discovery rates. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 64 0 (3): 0 479--498, 2002

  15. [23]

    Sun and T

    W. Sun and T. T. Cai. Oracle and adaptive compound decision rules for false discovery rate control. Journal of the American Statistical Association, 102 0 (479): 0 901--912, 2007

  16. [24]

    Sun and A

    W. Sun and A. C. McLain. Multiple testing of composite null hypotheses in heteroscedastic models. Journal of the American Statistical Association, 107 0 (498): 0 673--687, 2012

  17. [25]

    Sur and E

    P. Sur and E. J. Cand \`e s. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences, 116 0 (29): 0 14516--14525, 2019

  18. [26]

    Weinstein, Z

    A. Weinstein, Z. Ma, L. D. Brown, and C.-H. Zhang. Group-linear empirical bayes estimates for a heteroscedastic normal mean. Journal of the American Statistical Association, 113 0 (522): 0 698--710, 2018

  19. [27]

    Yekutieli

    D. Yekutieli. Adjusted bayesian inference for selected parameters. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 74 0 (3): 0 515--541, 2012

  20. [28]

    C.-H. Zhang. Empirical bayes and compound estimation of normal means. Statistica Sinica, 7 0 (1): 0 181--193, 1997

  21. [29]

    C.-H. Zhang. Compound decision theory and empirical bayes methods: invited paper. The Annals of Statistics, 31 0 (2): 0 379--390, 2003

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.