REVIEW 3 major objections 4 minor 7 references
Bayesian preposterior analysis yields sample allocations for heteroscedastic surveys that match or beat classical design-based and model-assisted alternatives.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:43 UTC pith:GQ2SXK5O
load-bearing objection A solid, honest extension of Bayesian optimal allocation to heteroscedastic ratio models; the clean derivation and broad simulation are the real contributions, but the 'known b' assumption is load-bearing and deserves a hard look in revision. the 3 major comments →
Bayesian Optimal Sample Design for Surveys with Heteroscedasticity
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that, under a stratified heteroscedastic normal model Y_hi ~ N(alpha_h X_hi, v_h X_hi^b), minimizing the approximate expected posterior variance of the finite population mean — Equation (13) — recovers the Bayesian optimal allocation as a function of known population X-values, one pilot sample, and the assumed b. The derivation breaks the expected posterior variance into a quadratic-form term and a design-dependent term, uses a chi-square result to evaluate the former and a first-order Taylor series for the latter. In simulations across 90 synthetic populations and in a public-charity revenue application, the resulting Bayesian allocation had lower root mean square error
What carries the argument
Equation (13): E_{D2|D1} Var(mean|D2,e1,b) ≈ sum_h EP(v_h)/N^2 · (n_h−1)/(n_h−3) · (N_h−n_h)[Zbar_h + (n_h/N_h S^2_xh + (N_h−n_h) Xbar_h^2)/(n_h Qbar_h^2(1+CV^2_qh))]. This is the preposterior expected posterior variance under squared error loss; it splits into a stratum-specific variance estimate from the pilot and the expected design variance of the auxiliary information, making the objective a smooth function of allocations that standard convex optimizers can minimize.
Load-bearing premise
The derivation treats the heteroscedasticity exponent b as known; the paper's own sensitivity analysis shows that assuming b=2 when the true b=0 in a highly skewed population raises RMSE by 240–281%.
What would settle it
Use the Bayesian allocation with b0=2 on the paper's highly skewed MOS1 populations with true b=0; if the relative RMSE increase versus the correctly specified allocation does not reach the reported 240–281%, the central optimality claim would be called into question.
If this is right
- Practitioners can allocate stratified samples without plugging in estimated strata variances; the pilot data are incorporated through the posterior expectation of the variance term.
- The allocation reduces to familiar Neyman-like forms when the heteroscedasticity exponent b is known and the model simplifies, but it accounts for heteroscedasticity explicitly through z_hi = x_hi^b.
- In the charity log-revenue application, the Bayesian allocation achieved relative RMSE 1.000 versus 1.427 for the classical design-based allocation and 1.014 for the plug-in ratio allocation.
- The Bayesian allocation produced more stable allocations across repeated pilot samples than plug-in methods, especially for highly skewed X and b=2.
- Misspecification of b matters mainly when the true b is near 0 and X is highly skewed; otherwise allocations are robust.
- The MC-based alternative to the Taylor series approximation for E[k(s2h)] is precomputed and allows continuous optimization with negligible loss of accuracy.
Where Pith is reading between the lines
- The same preposterior framework could extend to multivariate outcomes or non-normal errors (e.g., lognormal), where closed forms may not exist but the provided Monte Carlo emulation of E[k(s2h)] can be reused.
- The explicit separation of the quadratic form q_bh and the design term k(s2h) suggests that design optimality depends only on auxiliary-variable geometry, not on the outcome model's slope parameters; this may permit designs that are robust to slope misspecification.
- Since b is typically unknown, one could place a prior on b and integrate it out of Equation (13), directly addressing the paper's stated limitation about assuming b known.
- A testable extension: the MC-based E[k(s2h)] should be more reliable than the Taylor series for very small strata sample sizes (near the n_h≥4 boundary), even though both gave nearly identical results in the paper's simulations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian decision-theoretic approach to optimal sample allocation in stratified sampling under a heteroscedastic normal model Y_hi ~ N(alpha_h X_hi, v_h X_hi^b), with b assumed known and (alpha_h, v_h) unknown with a diffuse prior. Using Ericson's results, the authors derive an approximate closed form for the expected posterior variance of the finite population mean under squared error loss, given in Eq. (13), and optimize it subject to constraints via nonlinear programming. The proposed B-P strategy (Bayesian allocation and prediction estimator) is compared by simulation with Neyman-HT, Cochran's plug-in separate-ratio estimator (C-SR), and rule-of-thumb allocations across 90 artificial populations, and in an application to NCCS charity revenue data. The paper finds that B-P generally has lower or comparable RMSE than the alternatives, substantially so versus N-HT, and that misspecification of the heteroscedasticity exponent b can severely degrade B-P performance in some highly skewed populations.
Significance. If the result holds, the paper makes a useful contribution by extending the largely dormant Bayesian optimal sample design literature to heteroscedastic error structures, a setting relevant to establishment surveys. The derivation in Appendices A–B is algebraically coherent, and the Monte Carlo check described in §4.4 provides supporting evidence that the Taylor-series approximation to E[k(s_2h)] is accurate in the simulated settings. The paper is also transparent about its limitations: §7.2 explicitly acknowledges the known-b assumption and the lack of a full unknown-b pipeline. However, the headline claim that the proposed method 'do[es] as well or better' is not fully supported by the paper's own tables, and the practical utility of the method depends on how b is handled when it is unknown. The strengths are the clean derivation, the thorough simulation study with 90 populations, the real-data application, and the explicit sensitivity analysis.
major comments (3)
- [§3.2, §3.5, Eq. (13); §5.3, Table 6; §6.1] The optimality of B-P is derived under the assumption that b is known. The authors' own sensitivity analysis shows that when b is misspecified, B-P can be badly degraded: in Table 6, for MOS1 populations with true b=0, assuming b=2 increases RMSE by 240–281%. In practice b must be estimated, typically from the pilot sample of size m_h=15 per stratum, not from the full population. The NCCS application estimates b from the full data (posterior mean 0.55) and then treats it as true, avoiding the realistic two-stage uncertainty. Because Eq. (13) is misspecified when b is wrong, the allocation is no longer Bayes-optimal, and the 'do as well or better' claim is not supported for the realistic case of unknown b. I request a simulation of the full pipeline: estimate b from the pilot sample (e.g., MCMC or a simple estimator), plug the estimate into Eq. (13), and evaluate RMSE relative to C-SR and
- [Abstract; §5.1, Table 3] The claim that B-P 'do[es] as well or better' than the alternatives is contradicted by several cells in Table 3, where negative relative increases indicate that C-SR has lower RMSE than B-P. Examples include MOS2 baseline b=1 (-1.3%), MOS2 scenario 4 b=1 (-2.3%), MOS3 baseline b=1 (-3.4%), MOS3 scenario 5 b=1 (-3.6%), and MOS3 baseline b=0 (-0.8%), among others. The differences are small, but the wording is not literally accurate. The paper should either revise the claim to 'generally comparable or better, with small losses in some scenarios' or explain why those negative cells should be discounted.
- [§5.3 and §7.2] The sensitivity analysis for misspecified b is a useful step, but it does not address the practical problem of estimating b. It fixes b0 at incorrect values rather than simulating estimation of b from pilot data. The authors acknowledge in §7.2 that 'if allocations are sensitive to misspecified b, then performance of our proposed allocation methods will be unclear.' That is exactly the situation in the MOS1, b=0 case. The paper should either provide an additional analysis that treats b as unknown (e.g., integrate over the posterior of b, or use a pilot estimate with its uncertainty) or clearly scope the paper's contribution to the known-b setting and soften the practical recommendations.
minor comments (4)
- [§2.3, Step 5] Typo: 'onsidering' should be 'considering'.
- [§1, first paragraph] Typo: 'heterogenity' should be 'heterogeneity'.
- [Figure 1 note (Appendix C.2)] The phrase 'weights of by1/X^0.55' appears to be missing a space; should read 'weights of b = 1/X^0.55' or similar. This makes the sentence difficult to parse.
- [§4.5] The definition of Monte Carlo error via sd(\hat\theta) is standard, but the connection to the RMSE estimator used in Tables 2–6 could be stated more explicitly. It is not clear whether the reported Monte Carlo errors were included in the tables or only in the text.
Circularity Check
No significant circularity: Eq. (13) is derived from the assumed model and external Ericson results, not from the quantities it is used to choose.
full rationale
The central derivation chain is Eqs. (10)-(13): a stratified heteroscedastic normal model is assumed, a diffuse prior is placed on (alpha_h, v_h), Ericson's posterior results (external to the authors) give the posterior mean and variance, and a preposterior expectation over D2|D1 yields Eq. (13). The only estimated quantity entering Eq. (13) is EP(v_h), computed from the pilot D1 via Equation (57); this is the intended mechanism by which pilot data influence the allocation, and the allocation is then evaluated on new main-study samples. Nothing in Eq. (13) is defined in terms of the allocation it is used to choose, and no fitted value is renamed as a prediction. The paper's own sensitivity analysis (Section 5.3, Table 6) and limitation statement (Section 7.2) explicitly concede that b is assumed known and that misspecification can degrade performance; this is a scope/robustness limitation, not a circular derivation. The comparisons against N-HT, C-SR, and RT allocations are simulation evaluations of independently defined estimators; that the artificial populations are generated from the same model is an external-validity caveat acknowledged in Section 7.2, not evidence that the derivation merely restates its inputs. Self-citations (Coffey et al. 2020; Wagner et al. 2020; West et al. 2021) appear only in background discussion of priors for nonresponse follow-up, not as load-bearing support for the main allocation result.
Axiom & Free-Parameter Ledger
free parameters (1)
- b (heteroscedasticity exponent) =
0.55 in log-revenue NCCS application; 1.19 in original-scale application; simulated at {0, 0.5, 1, 1.5, 2}
axioms (5)
- domain assumption Diffuse prior π(α_h, 1/v_h) ∝ v_h yields a normal-gamma posterior for (α_h, v_h).
- domain assumption Within each stratum, Y_{hi} | (X, α_h, v_h, b) ~ N(α_h X_{hi}, v_h X_{hi}^b) with known X_{hi} and known b.
- domain assumption Stratified simple random sampling with independent pilot and main phases, and fixed strata with known population auxiliary totals.
- domain assumption First-order Taylor approximation E[num/den] ≈ E[num]/E[den] is accurate for large n_h.
- standard math Ericson's (1969a) posterior mean and variance results for a finite population mean under a normal-gamma model are correct and applicable.
read the original abstract
We develop a Bayesian optimal sample allocation approach for stratified sampling in heteroscedastic populations. Existing optimal allocation theory typically assumes knowledge of certain design parameters (e.g., strata variances) that may be unknown, leading practitioners to substitute in survey-based estimates when planning samples, often without considering the effects of this substitution on sample efficiency. Bayesian decision theory for optimal experimental design avoids such substitutions and can be applied to sample allocation. Bayesian sample optimization methods were studied heavily from the mid-1960s through early 1980s, but have been overlooked since, in spite of modern computing advances that have facilitated a proliferation of Bayesian methods in other areas of statistics. A limitation of this early Bayesian sample design work is that it did not accommodate heteroscedastic error structures, which underlie commonly used ratio estimation models. Our paper, which optimizes the design under a univariate regression model with heteroscedastic errors, addresses this limitation of earlier work, while illustrating the Bayesian approach to design. We identify the optimal Bayesian allocation under our model, then compare performance of the proposed Bayesian sampling strategy with that of key design-based and model-assisted alternatives across several settings, finding that the proposed methods do as well or better than the alternatives under the scenarios considered. We apply our methods in analyzing revenues of public charities, using publicly available IRS Form 990 data.
Figures
Reference graph
Works this paper leans on
-
[5]
corrs, dec
Inc. corrs, dec. slopes b = 0 -0.2 0.3 -0.2 0.6 0.3 0.5 0.1 -0.3 -0.2 -0.3 -0.4 -0.2 0.7 -0.2 -0.1 0.2 0.0 -0.4 b = 0.5 0.7 0.5 0.1 0.6 -0.9 -0.0 0.1 0.1 0.2 0.1 0.3 0.4 -0.4 0.2 -0.1 0.3 -0.2 0.3 b = 1 0.0 -0.3 0.6 1.0 -0.8 -0.3 -0.0 0.1 -0.1 -0.5 0.1 -0.6 -0.2 -0.4 -0.0 -0.1 -0.1 -0.4 b = 1.5 -0.6 0.4 0.4 1.1 0.6 0.7 0.2 0.1 -0.1 -0.3 0.1 0.1 -0.2 -0.0 ...
-
[9]
corrs, dec
Inc. corrs, dec. slopes b = 0 94.9 94.0 94.9 95.4 94.6 94.0 95.3 93.8 95.6 95.2 95.4 95.1 94.3 95.0 94.8 95.3 94.8 94.0 b = 0.5 95.4 95.7 95.0 94.7 95.8 95.2 95.5 94.8 95.3 94.4 95.2 94.4 94.6 95.0 93.9 95.7 96.0 95.4 b = 1 94.8 94.8 95.8 95.8 95.5 95.4 94.8 95.1 95.9 94.7 93.9 95.3 94.4 94.8 94.7 96.4 95.3 93.6 b = 1.5 95.4 95.1 94.2 94.9 95.4 94.9 95.2 ...
-
[12]
corrs, dec
Inc. corrs, dec. slopes b = 0 -0.6 -1.1 0.1 -0.0 -0.4 -0.6 -0.0 -0.1 0.0 -0.0 -0.0 0.4 0.2 0.1 -0.1 0.4 0.1 -0.2 b = 1 0.0 0.2 0.2 -0.2 -0.6 0.1 0.3 0.8 -0.0 -0.1 -0.0 -0.4 0.6 0.1 -0.0 0.4 -0.2 -0.1RelBias*1000 b = 2 -0.4 -0.1 0.2 -0.7 -0.9 1.7 0.0 0.3 0.0 0.5 -0.2 1.2 -0.1 0.2 0.0 0.4 -0.0 0.0 b = 0 93.8 95.4 96.1 95.2 94.0 95.3 95.1 95.4 95.0 95.3 95.3...
-
[14]
corrs, dec
Inc. corrs, dec. slopes
-
[15]
corrs, dec
Inc. corrs, dec. slopes b0 = 0 1.0 1.1 0.4 -1.5 0.2 2.4 -0.0 0.6 0.1 0.2 0.0 -0.2 -0.7 0.6 0.5 2.7 1.9 2.3 b0 = 0.5 0.9 -0.6 0.5 -1.6 0.5 0.8 -0.6 0.1 0.2 0.4 0.9 -0.8 -0.5 2.3 0.7 3.1 1.7 2.0 b0 = 1 1.0 -0.7 0.3 -1.3 -0.3 2.0 -0.2 -0.4 -0.1 0.0 0.0 1.1 -0.4 1.3 0.4 2.9 2.5 2.6 b0 = 1.5 2.2 -1.8 0.1 -3.4 0.4 -0.1 0.0 0.2 -0.1 -0.1 -0.0 -0.2 -0.6 1.5 0.4 3...
2008
-
[187]
Hansen Lecture 2022: The Evolution of the Use of Models in Survey Sampling
DOI: https://doi.org/10.1093/jssam/smv038. Valliant, R. 2023. “Hansen Lecture 2022: The Evolution of the Use of Models in Survey Sampling.”Journal of Survey Statistics and Methodology12 (2): 275–304. DOI: https: //doi.org/10.1093/jssam/smad021. Valliant, R., J. A. Dever, and F. Kreuter. 2013.Practical Tools for Designing and Weighting Survey Samples.Sprin...
arXiv 2023
-
[357]
Bayesian Inference in Finite Populations
Wiley. . 1972.Subjective Bayesian Models in Sampling Finite Populations: Stratification. Technical report 15. Ann Arbor, Michigan: University of Michigan, Department of Statistics. . 1973.A Bayesian Approach to Two-Stage Sampling.Technical report 26. Ann Arbor, Michigan: University of Michigan, Department of Statistics. . 1975.A Bayesian Approach to Two-S...
arXiv 1972
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.