{"id":"c3b3b31c-e138-4ad5-9c42-79bf4cc4b595","arxiv_id":"2607.21775","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A Bayesian decision-theoretic allocation rule is derived for stratified samples under a heteroscedastic ratio model, and it matches or improves on standard alternatives in most simulations and a charity revenue application.","lead":"This paper derives a Bayesian rule for spreading a survey sample across strata when within-stratum variability grows with a known size measure. In simulations and an IRS-based application, the rule matches or beats standard Neyman and ratio-estimator allocations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"B-P's optimality holds only for known b; with b estimated from pilot data, RMSE gains may vanish — the paper's own Table 6 shows 240–281% RMSE inflation under misspecification.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: b is treated as known in the derivation and implementation, but must be supplied in practice. This is the most load-bearing condition because Eq. (13) is only the expected posterior loss when b is correct; if b is misspecified, the minimized quantity does not correspond to the actual expected posterior loss, so the allocation is not Bayes-optimal. The paper's own sensitivity analysis (Table 6) demonstrates that the method can degrade severely in skewed populations, losing its advantage over simpler strategies. Although the paper acknowledges this limitation and provides sensitivity analysis, it does not test the realistic scenario where b is estimated from the pilot sample, which is the actual operational condition. The proposed test would settle whether the method retains its claimed advantages when b is unknown. The reader's CONDITIONAL verdict remains appropriate: the analytical core is sound under the stated assumptions, but external validity is limited by the known-b requirement, and the missing pilot-based b-estimation simulation is a concrete gap.","tokens_in":38726,"tokens_out":14002,"duration_ms":141205,"concrete_test":"Use the simulated MOS1 true b=0 and b=1 populations (§5.3). For each of 1000 pilot samples, estimate b from the pilot data (e.g., via MCMC under the assumed model, as in §6.1, or via a profile-likelihood method), then run the full B-P pipeline with the estimated b for both allocation and inference. Compare RMSE to (a) B-P with true b and (b) C-SR. If RMSE(estimated-b B-P) exceeds RMSE(C-SR), or if the RMSE increase over correct-b B-P approaches the 240–281% range in Table 6, the practical claim of superiority fails when b is unknown. As a complementary check, in the NCCS application draw b from its MCMC posterior for each simulated pilot and use that draw in Eq. (13); compare RMSE to the fixed-b=0.55 analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that B-P allocations 'do as well or better' depends on Eq. (13), which conditions on b as known (§3.2, §3.5). In practice b must be estimated; §5.3 shows that for highly skewed MOS1 populations with true b=0, assuming b=2 increases RMSE by 240–281% (Table 6). The NCCS application (§6.1) estimates b from the full population (posterior mean 0.55, 95% CI [0.25, 0.66]) and then treats it as known; in a real rollout only the pilot sample would be available (n=360 in the application, m_h=15 in the simulations), and a pilot-based b estimate would have substantially larger uncertainty. The paper does not simulate the realistic pipeline of estimating b from the pilot and then using it in Eq. (13). Because the objective function itself is misspecified when b is wrong, the allocation is no longer Bayes-optimal, so the 'do as well or better' claim is not supported for the realistic case of unknown b.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Bayesian decision-theoretic approach to optimal sample allocation in stratified sampling under a heteroscedastic normal model Y_hi ~ N(alpha_h X_hi, v_h X_hi^b), with b assumed known and (alpha_h, v_h) unknown with a diffuse prior. Using Ericson's results, the authors derive an approximate closed form for the expected posterior variance of the finite population mean under squared error loss, given in Eq. (13), and optimize it subject to constraints via nonlinear programming. The proposed B-P strategy (Bayesian allocation and prediction estimator) is compared by simulation with Neyman-HT, Cochran's plug-in separate-ratio estimator (C-SR), and rule-of-thumb allocations across 90 artificial populations, and in an application to NCCS charity revenue data. The paper finds that B-P generally has lower or comparable RMSE than the alternatives, substantially so versus N-HT, and that misspecification of the heteroscedasticity exponent b can severely degrade B-P performance in some highly skewed populations.","tokens_in":39009,"tokens_out":4223,"duration_ms":46775,"significance":"If the result holds, the paper makes a useful contribution by extending the largely dormant Bayesian optimal sample design literature to heteroscedastic error structures, a setting relevant to establishment surveys. The derivation in Appendices A–B is algebraically coherent, and the Monte Carlo check described in §4.4 provides supporting evidence that the Taylor-series approximation to E[k(s_2h)] is accurate in the simulated settings. The paper is also transparent about its limitations: §7.2 explicitly acknowledges the known-b assumption and the lack of a full unknown-b pipeline. However, the headline claim that the proposed method 'do[es] as well or better' is not fully supported by the paper's own tables, and the practical utility of the method depends on how b is handled when it is unknown. The strengths are the clean derivation, the thorough simulation study with 90 populations, the real-data application, and the explicit sensitivity analysis.","major_comments":[{"comment":"The optimality of B-P is derived under the assumption that b is known. The authors' own sensitivity analysis shows that when b is misspecified, B-P can be badly degraded: in Table 6, for MOS1 populations with true b=0, assuming b=2 increases RMSE by 240–281%. In practice b must be estimated, typically from the pilot sample of size m_h=15 per stratum, not from the full population. The NCCS application estimates b from the full data (posterior mean 0.55) and then treats it as true, avoiding the realistic two-stage uncertainty. Because Eq. (13) is misspecified when b is wrong, the allocation is no longer Bayes-optimal, and the 'do as well or better' claim is not supported for the realistic case of unknown b. I request a simulation of the full pipeline: estimate b from the pilot sample (e.g., MCMC or a simple estimator), plug the estimate into Eq. (13), and evaluate RMSE relative to C-SR and","section":"§3.2, §3.5, Eq. (13); §5.3, Table 6; §6.1"},{"comment":"The claim that B-P 'do[es] as well or better' than the alternatives is contradicted by several cells in Table 3, where negative relative increases indicate that C-SR has lower RMSE than B-P. Examples include MOS2 baseline b=1 (-1.3%), MOS2 scenario 4 b=1 (-2.3%), MOS3 baseline b=1 (-3.4%), MOS3 scenario 5 b=1 (-3.6%), and MOS3 baseline b=0 (-0.8%), among others. The differences are small, but the wording is not literally accurate. The paper should either revise the claim to 'generally comparable or better, with small losses in some scenarios' or explain why those negative cells should be discounted.","section":"Abstract; §5.1, Table 3"},{"comment":"The sensitivity analysis for misspecified b is a useful step, but it does not address the practical problem of estimating b. It fixes b0 at incorrect values rather than simulating estimation of b from pilot data. The authors acknowledge in §7.2 that 'if allocations are sensitive to misspecified b, then performance of our proposed allocation methods will be unclear.' That is exactly the situation in the MOS1, b=0 case. The paper should either provide an additional analysis that treats b as unknown (e.g., integrate over the posterior of b, or use a pilot estimate with its uncertainty) or clearly scope the paper's contribution to the known-b setting and soften the practical recommendations.","section":"§5.3 and §7.2"}],"minor_comments":[{"comment":"Typo: 'onsidering' should be 'considering'.","section":"§2.3, Step 5"},{"comment":"Typo: 'heterogenity' should be 'heterogeneity'.","section":"§1, first paragraph"},{"comment":"The phrase 'weights of by1/X^0.55' appears to be missing a space; should read 'weights of b = 1/X^0.55' or similar. This makes the sentence difficult to parse.","section":"Figure 1 note (Appendix C.2)"},{"comment":"The definition of Monte Carlo error via sd(\\hat\\theta) is standard, but the connection to the RMSE estimator used in Tables 2–6 could be stated more explicitly. It is not clear whether the reported Monte Carlo errors were included in the tables or only in the text.","section":"§4.5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the analytical core appears sound. My main concern is the practical handling of the heteroscedasticity exponent b: the paper's own sensitivity analysis shows that the method can be highly sensitive to b in certain populations, yet the realistic case of pilot-based estimation of b is not simulated. This is fixable and should be addressed before publication. If the authors add the requested unknown-b simulation and revise the 'do as well or better' claim, I would support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first real extension of the old Bayesian optimal allocation literature—Draper & Guttman, Ericson, Rao & Ghangurde—to heteroscedastic ratio models of the form v_h X^b. The main analytical result, the approximate expected posterior variance in Eq. (13), is genuine work and the derivation in Appendices A–B holds up. Second, the headline claim \"as well or better\" is conditional on knowing b. The stress-test concern about estimating b from pilot data is on target, and the paper itself flags the assumption in Section 7.2. That is not a hidden flaw, but it is the soft spot.\n\nWhat the paper does well: the algebra is coherent, the Monte Carlo check supports the Taylor approximation, and the simulation study is broad—90 populations, three skewness levels, five b values, six correlation/slope scenarios. Comparisons against Neyman-HT and Cochran-SR are fair, and the NCCS application is a useful real-data test. The paper is also honest that C-SR is competitive there (relative RMSE 1.014 vs 1.000), so it does not oversell.\n\nWhere it is soft: the b-known assumption is load-bearing. The paper estimates b from the same NCCS data and then treats it as true in the allocation. Table 6 shows that for highly skewed MOS1 populations with true b=0, assuming b=2 inflates RMSE by 240–281%. The paper does not simulate the realistic pipeline of estimating b from a small pilot and plugging it into Eq. (13). That is a genuine gap in the empirical case, not a manufactured one. The simulations are also generated from the assumed model, which limits external validity. Minor points: no code or data are shipped, and Monte Carlo error for the RMSE comparisons is computed but not shown in the tables. These are worth fixing but not deal-breakers.\n\nWho this is for: survey methodologists, especially at statistical agencies dealing with establishment surveys. A serious referee should engage with it. My recommendation: send it out. The revision should add a pilot-estimate-b simulation and ideally ship code, but the core derivation is sound and the paper fills a real gap in the literature.","headline":"A solid, honest extension of Bayesian optimal allocation to heteroscedastic ratio models; the clean derivation and broad simulation are the real contributions, but the 'known b' assumption is load-bearing and deserves a hard look in revision.","tokens_in":39472,"tokens_out":2167,"would_cite":true,"duration_ms":24182,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian preposterior analysis yields sample allocations for heteroscedastic surveys that match or beat classical design-based and model-assisted alternatives.","keywords":["sample allocation","stratified sampling","heteroscedasticity","Bayesian decision theory","preposterior analysis","finite population mean","ratio estimation","expected posterior variance"],"falsifier":"Use the Bayesian allocation with b0=2 on the paper's highly skewed MOS1 populations with true b=0; if the relative RMSE increase versus the correctly specified allocation does not reach the reported 240–281%, the central optimality claim would be called into question.","tokens_in":38629,"feed_emoji":"📊","tokens_out":4559,"duration_ms":45158,"temperature":0.7,"pith_summary":"The paper tries to establish that Bayesian decision theory can produce sample allocations for stratified surveys that account both for heteroscedasticity and for uncertainty in design parameters, and that the resulting allocations are at least as accurate as standard design-based and model-assisted alternatives. It derives an approximate closed-form expected posterior variance for the finite population mean under a heteroscedastic normal regression model, then minimizes it by convex optimization. In 90 simulated populations and in a real charity-revenue application, the Bayesian allocation matches or outperforms the classical optimal allocation and the plug-in separate ratio allocation, and it leads to more stable allocations across repeated pilot samples.","feed_headline":"Bayesian sample allocation beats classical designs in heteroscedastic surveys","feed_subtitle":"Uses pilot data to minimize expected posterior variance; matches or beats plug-in alternatives in simulations and charity data.","key_machinery":"Equation (13): E_{D2|D1} Var(mean|D2,e1,b) ≈ sum_h EP(v_h)/N^2 · (n_h−1)/(n_h−3) · (N_h−n_h)[Zbar_h + (n_h/N_h S^2_xh + (N_h−n_h) Xbar_h^2)/(n_h Qbar_h^2(1+CV^2_qh))]. This is the preposterior expected posterior variance under squared error loss; it splits into a stratum-specific variance estimate from the pilot and the expected design variance of the auxiliary information, making the objective a smooth function of allocations that standard convex optimizers can minimize.","core_discovery":"The central claim is that, under a stratified heteroscedastic normal model Y_hi ~ N(alpha_h X_hi, v_h X_hi^b), minimizing the approximate expected posterior variance of the finite population mean — Equation (13) — recovers the Bayesian optimal allocation as a function of known population X-values, one pilot sample, and the assumed b. The derivation breaks the expected posterior variance into a quadratic-form term and a design-dependent term, uses a chi-square result to evaluate the former and a first-order Taylor series for the latter. In simulations across 90 synthetic populations and in a public-charity revenue application, the resulting Bayesian allocation had lower root mean square error","pith_inferences":["The same preposterior framework could extend to multivariate outcomes or non-normal errors (e.g., lognormal), where closed forms may not exist but the provided Monte Carlo emulation of E[k(s2h)] can be reused.","The explicit separation of the quadratic form q_bh and the design term k(s2h) suggests that design optimality depends only on auxiliary-variable geometry, not on the outcome model's slope parameters; this may permit designs that are robust to slope misspecification.","Since b is typically unknown, one could place a prior on b and integrate it out of Equation (13), directly addressing the paper's stated limitation about assuming b known.","A testable extension: the MC-based E[k(s2h)] should be more reliable than the Taylor series for very small strata sample sizes (near the n_h≥4 boundary), even though both gave nearly identical results in the paper's simulations."],"forward_implications":["Practitioners can allocate stratified samples without plugging in estimated strata variances; the pilot data are incorporated through the posterior expectation of the variance term.","The allocation reduces to familiar Neyman-like forms when the heteroscedasticity exponent b is known and the model simplifies, but it accounts for heteroscedasticity explicitly through z_hi = x_hi^b.","In the charity log-revenue application, the Bayesian allocation achieved relative RMSE 1.000 versus 1.427 for the classical design-based allocation and 1.014 for the plug-in ratio allocation.","The Bayesian allocation produced more stable allocations across repeated pilot samples than plug-in methods, especially for highly skewed X and b=2.","Misspecification of b matters mainly when the true b is near 0 and X is highly skewed; otherwise allocations are robust.","The MC-based alternative to the Taylor series approximation for E[k(s2h)] is precomputed and allows continuous optimization with negligible loss of accuracy."],"fun_headline_variants":["Bayesian sample allocation matches or beats plug-in in heteroscedastic surveys","Bayesian sampling for heteroscedastic strata: matches or beats classical","Heteroscedastic surveys: Bayesian allocation is as good or better"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The derivation treats the heteroscedasticity exponent b as known; the paper's own sensitivity analysis shows that assuming b=2 when the true b=0 in a highly skewed population raises RMSE by 240–281%.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian sample allocation matches or beats plug-in in heteroscedastic surveys","Bayesian sampling for heteroscedastic strata: matches or beats classical","Heteroscedastic surveys: Bayesian allocation is as good or better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000799,"raw_usage":{"total_tokens":3362,"prompt_tokens":766,"completion_tokens":2596,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":2541}},"tokens_in":510,"tokens_out":2596,"duration_ms":21131,"temperature":1.0,"reasoning_tokens":2541,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:43:15.688437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the Bayesian allocation with b0=2 on the paper's highly skewed MOS1 populations with true b=0; if the relative RMSE increase versus the correctly specified allocation does not reach the reported 240–281%, the central optimality claim would be called into question.","supporting_citations":[],"review_version":1}