{"id":"81413ae7-a696-40cc-ab8f-9ba538d6c277","arxiv_id":"2411.18549","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Plug-in variance estimators for Bowley's b2 and Groeneveld-Meeden b3 are proposed for finite population sampling, with simulation evidence of adequate confidence interval coverage.","lead":"This paper develops standard errors and confidence intervals for two skewness measures when computed from survey samples of a finite population. Simulations suggest the intervals usually have at least the nominal coverage, but the theory behind them is explicitly heuristic.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The normal-CI claim depends on unproven Assumptions B1 and B2 (Appendix A): the von Mises expansion with negligible remainder and the CLT for the linearized integral are asserted rather than established for the step-function Hajek and calibration cdf estimators, and the paper explicitly declines…","rationale":"The reader's verdict (CONDITIONAL) is appropriate. The paper's central claim—asymptotic normality of plug-in skewness estimators with the given variance and valid normal CIs—is explicitly conditioned on B1 and B2, which are not proven. My stress-test confirms this is the single most load-bearing gap: if B1 or B2 fails, the asymptotic variance formula (6) may be wrong or the CLT may not hold, and the CI coverage can be mis-calibrated. The paper's heuristic Appendix A requires condition A (continuous positive densities for F and Fhat), which fails for the step-function estimators; the remainder R_N needs a substantive uniform bound that is not supplied. The paper's own limitations (no proof of B1/B2, no analysis of the density estimator bias) are candidly acknowledged, which supports a conditional acceptance rather than rejection: the simulations provide some evidence of adequate coverage, and the methodology is a reasonable extension of existing plug-in survey-sampling techniques. I recommend no change to the reader's verdict (UNCHANGED). A concrete test—an independent analytical bound on R_N or a simulation with a small density at the median—would either confirm or refute the conjecture.","tokens_in":13327,"tokens_out":13027,"duration_ms":120277,"concrete_test":"Independently re-derive Assumption B1 for the Hajek estimator: write the exact remainder R_N in eq. (4) for bb3 and show whether it is o_p(V_N) when Fhat is the Hajek step-function cdf. Specifically, check if the second-order term contains a factor (Fhat(nu)-F(nu))^2 / f(nu) plus a term involving the density estimator (3); if the density estimator's bias contributes a non-vanishing term to R_N, B1 fails. A numerical counterpart: simulate N=800, n=40/80 from a population with f(median)=0.1 (e.g., a mixture of two normals with a narrow gap) and compute the 95% CI coverage for b3; if coverage drops below 0.93, the 'broad conditions' do not cover small but positive f(nu), undermining the claimed validity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"B1 and B2 are the load-bearing assumptions: they convert the heuristic functional delta-method into a rigorous CLT. The paper presents no proof and only states a conjecture with references. This is not an internal inconsistency, but it is a missing foundation: the finite-population cdf estimators are step functions, so the formal derivation in Appendix A, which requires continuous positive densities for F and Fhat at every lambda (condition A), cannot be applied directly. The remainder R_N in eq. (4) contains second-order terms in the empirical process at the quantiles, divided by density factors; controlling these terms requires a uniform Bahadur-type representation for the sample quantiles and for delta, and a rate condition on the density estimator. If B1 fails, the reported variance V_N may not be the asymptotic variance of bb, and the normal intervals need not have the claimed coverage. The simulations in Appendix B are encouraging but do not settle the question: only two populations and two designs are used, and the variance estimators' relative bias is often above 0.3 at n=80, which is consistent with slow or failing convergence. The paper's own statements ('we do not investigate sufficient conditions', 'a formal analysis of this bias lies outside of the scope') confirm the gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops plug-in estimators for two quantile-based skewness measures, Bowley's b2(r) and the Groeneveld-Meeden b3, in finite population sampling. Estimators are formed by inserting the Hajek or a Deville-Sarndal calibration cdf estimator into the definitions of the skewness measures. The paper derives asymptotic variance formulae by a functional delta method argument, proposes Horvitz-Thompson and Sen-Yates-Grundy variance estimators, and evaluates the resulting normal confidence intervals in a simulation study with two synthetic populations, two sampling designs, and two sample sizes. The central claim is that, under broad but unproved conditions, the standardized plug-in estimators are asymptotically normal with variance given by the proposed formulae, so that the proposed intervals have nominal coverage asymptotically.","tokens_in":13616,"tokens_out":2768,"duration_ms":27090,"significance":"If the central asymptotic claim holds, the paper makes a useful contribution by extending quantile-based skewness inference from i.i.d. settings to design-based finite population sampling and by providing explicit influence-function-type variance formulae for two cdf estimators. The paper is commendably honest: it labels the functional delta method argument as heuristic, explicitly states Assumptions B1 and B2 as conjectures rather than theorems, and openly notes that the density estimator bias is not formally analyzed. The simulation study, though limited, provides encouraging evidence that coverage rates are often conservative and that variance estimator stability is comparable to that reported in earlier survey-sampling work. The main limitation is that the entire normal-CI procedure rests on unproved high-level assumptions, so the paper currently supplies methodology and evidence but not a rigorous foundation for its headline inferential claim.","major_comments":[{"comment":"The paper's central claim that (bb - b)/V_N converges in distribution to N(0,1) is stated to follow from Assumptions B1 and B2, but these assumptions are only conjectured: the text says 'we do not investigate sufficient conditions under which B1 and B2 hold' and merely cites Conti and Marella (2015), Han and Wellner (2021), and Dey and Chaudhuri (2024). This is load-bearing because Condition A, under which the von Mises expansion is formally derived, requires continuous and positive densities for F(t) and Fhat(t), while the Hajek and calibration cdf estimators are step functions. The remainder R_N in equation (4) contains second-order terms at estimated quantiles divided by density factors, and controlling it for step-function cdf estimators requires a uniform Bahadur-type representation and rate conditions on the density estimator, none of which are stated or proved. Without a proof or at least a precise statement of sufficient conditions, the variance formula in equation (6) is not established as the asymptotic variance of bb, and the claimed validity of normal confidence intervals is unsupported.","section":"Appendix A, Assumptions B1 and B2"},{"comment":"The proposed variance estimators depend on the density estimator fhat(nu_r) defined in equation (3), and the paper acknowledges that this estimator carries bias whose formal analysis is outside the scope. The simulation results show that this bias is not negligible: in Table 3 the relative bias of the variance estimator for bb2,Ha(0.75) is 1.363 at n=40 and 0.370 at n=80, and several other entries exceed 0.30 at n=80 (Tables 3, 5, 7). These numbers are consistent with the paper's own conjecture that the bias decreases slowly, but they also mean that the variance estimators can severely overestimate the MSE, which is relevant to the practical claim that normal intervals are valid. The paper should either provide a first-order analysis of the density estimator bias and its effect on V-hat, or propose a bias-corrected density estimator, or clearly restrict the validity claim to settings where the bias is asymptotically negligible and show that the simulations support that restriction.","section":"Section 2, equation (3), and Appendix B"},{"comment":"The simulation study uses only two populations, both generated from the same model with normal errors and lognormal X, two sampling designs (SRS and stratified SRS), and two sample sizes (n=40,80) from a population of N=800. This is too narrow to substantiate the statement in the introduction that normal confidence intervals work 'under broad conditions.' In particular, the behavior of quantile-based skewness estimators and of the Woodruff-based density estimator may depend strongly on the local density at the relevant quantiles and on the design, and the simulations do not explore heavy-tailed distributions, unequal-probability designs without stratification, or populations where the target quantiles lie in regions of low density. At minimum, the conclusions should be tempered to describe the evidence as preliminary, and the simulation design should be expanded or justified as representative of the settings for which the theory is intended.","section":"Section 3 and Appendix B"}],"minor_comments":[{"comment":"There is a garbled sentence in the derivation of the von Mises derivative: 'In order to make sure that However, partial_nu_lambda/partial_lambda ...' appears to be an incomplete edit. This should be corrected.","section":"Appendix A, paragraph after equation (4)"},{"comment":"The notation fhat(nu_r) is introduced with square brackets that are easy to confuse with a floor or indicator function; a clearer notation such as widehat{f}(hat{nu}_r) would improve readability.","section":"Section 2, equation (3)"},{"comment":"The claim that the calibration equations (2) are 'always solvable' should be accompanied by a reference or a brief condition on the support of X and the sample design, since solvability of exponential calibration equations is not completely unconditional.","section":"Section 2, equation (2)"},{"comment":"The tables report only the Sen-Yates-Grundy variance estimators; the paper mentions Horvitz-Thompson-type estimators as alternatives, but does not report their performance. A sentence noting that SYG was chosen because of fixed-size designs is present, but it would be helpful to state explicitly that the HT versions were not evaluated.","section":"Appendix B"},{"comment":"The paper cites Groeneveld (1991) for influence functions of b2(r) and b3 but does not explain how those influence functions relate to the g-functions in Appendix A beyond a one-sentence remark; a short display of the connection would make the paper more self-contained.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and the proposed methodology is plausible, but the central asymptotic claim is explicitly conditional on unproved assumptions that are not merely technical details. Because the paper positions itself as providing finite population inference for skewness measures, I think the gap between the heuristic delta-method derivation and the stated CLT must be addressed before publication. The revision could succeed if the author proves or precisely states sufficient conditions for B1 and B2 for the Hajek and calibration cdf estimators, or alternatively reframes the paper as a simulation-based evaluation of a heuristic procedure. The latter would reduce the paper's scope but could still be publishable if the simulation study is expanded and the claims are adjusted accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper does something useful and is refreshingly honest about what it doesn't prove. It gives plug-in estimators and variance estimators for Bowley's b2 and Groeneveld-Meeden's b3 under finite population sampling, with the functional delta method as the engine. The influence functions were already in Groeneveld (1991), so the novelty is the finite-population variance machinery and the simulation evaluation, not the discovery of a new method. That's fine; it fills a real gap for survey practitioners who want standard errors for these descriptive measures.\n\nWhat's good: the derivation is clear, the variance estimators for the Hajek and calibration cdf estimators are carefully laid out, and the simulation study is honest. The author reports coverage rates, tail errors, relative bias and stability of the variance estimators, and they note when the variance estimator's relative bias is large (e.g., 1.36 in Table 3) and that it doesn't always shrink with n. That kind of transparency is worth a lot.\n\nThe soft spots are real but not hidden. The main theorem rests on Assumptions B1 and B2 in Appendix A, which are the von Mises expansion with negligible remainder and a CLT for the integral term. The paper explicitly says it does not investigate sufficient conditions and only conjectures they hold. That's a load-bearing gap: without a proof, the entire normal-CI claim is heuristic. The density estimator in (3) is acknowledged to carry unanalyzed bias, and the large relative biases in the variance estimator are consistent with that bias converging slowly. The simulations cover only two populations and two designs, so they're suggestive, not decisive.\n\nNone of this kills the paper. The author says the right things about what is and isn't established, and the empirical evidence is consistent with the claims. But a referee should push for either a proof of B1/B2 under finite-population asymptotics (there are tools in Conti-Marella, Han-Wellner, Dey-Chaudhuri) or a careful analysis of the density estimator's bias, and ideally code/data release.\n\nWho should read it: survey methodologists who need variance estimates for skewness measures, and anyone applying plug-in inference to distribution-based parameters. It deserves a serious referee, though the outcome should be conditional on the theoretical gaps being addressed or substantially reduced.","headline":"Honest, useful extension of plug-in inference to two quantile-based skewness measures, but the asymptotic theory is explicitly conjectural and the density estimator bias is unquantified; worth refereeing, not worth treating as settled.","tokens_in":14098,"tokens_out":2026,"would_cite":true,"duration_ms":19016,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper derives asymptotic normality for plug-in estimators of Bowley's skewness and the Groeneveld–Meeden index in finite population sampling.","keywords":["skewness","Bowley's index","Groeneveld–Meeden index","finite population inference","functional delta method","calibration estimator","Hájek estimator","Sen–Yates–Grundy variance estimator"],"falsifier":"Generate a finite population with a known skewness value and a design with strongly unequal inclusion probabilities, then compute $\\widehat{b}$ and the variance estimate over many samples and check the empirical distribution of $(\\widehat{b} - b)/\\widehat{V}$. If the 95% interval covers in well under 90% of samples, the central claim is false for that setting. A more direct check is to compute the remainder $R_N$ from the von Mises expansion for increasing $N$ and see whether it vanishes relative to the integral term.","tokens_in":13129,"feed_emoji":"📊","tokens_out":8031,"duration_ms":63253,"temperature":0.7,"pith_summary":"This paper tries to establish that two classic skewness measures—Bowley's $b_2(r)$ and the Groeneveld–Meeden $b_3$—can be estimated from a finite population sample with a variance estimator that makes normal confidence intervals valid. For plug-in estimators built from the Hájek cdf estimator and from a Deville–Särndal calibration cdf estimator, the paper derives asymptotic variance formulas by a von Mises (functional delta) expansion and proposes Horvitz–Thompson and Sen–Yates–Grundy type variance estimators. The paper's central claim is asymptotic normality of the standardized estimator, $(\\widehat{b} - b)/V \\to N(0,1)$, under broad conditions, plus the assertion that the proposed variance estimators approximate $V$. Simulations under simple and stratified random sampling show coverage rates at or above nominal levels for $b_2(0.75)$ and $b_3$, while the variance estimators retain a noticeable bias that the paper traces to plug-in density estimation.","feed_headline":"Skewness estimates from surveys get normal confidence intervals","feed_subtitle":"Bowley's and Groeneveld-Meeden skewness estimates get variance formulas and normal confidence intervals in simulations.","key_machinery":"The machinery is the von Mises expansion of the plug-in estimator, also called the functional delta method. For $b_3$ the paper builds a one-parameter path $F_\\lambda = F + \\lambda(\\widehat{F} - F)$ and differentiates the functional to obtain the influence weight $g(t)$; the derivative exists only under a smoothness condition, so the paper replaces it with assumption B1 that the expansion holds with a negligible remainder even for the discrete Hájek and calibration cdf estimators. The variance calculation then reduces to the variance of the weighted sample sum $\\sum_{i\\in s} d_i g(y_i)$, and the proposed variance estimators are the standard Horvitz–Thompson and Sen–Yates–Grundy forms applied to those weighted sums. A plug-in density estimator for $f(\\nu_r)$ completes the construction, and the paper notes that this density estimator is the suspected source of the variance estimators' bias.","core_discovery":"The claimed discovery is that quantile-based skewness measures, despite being nonlinear functionals of a step-function cdf estimator, admit the same linearization that makes survey inference work for smooth functionals. Specifically, the paper asserts that $\\widehat{b}_\\bullet - b_\\bullet = \\int g(t)\\,d[\\widehat{F}(t) - F(t)] + R_N$ with $R_N$ asymptotically negligible, where $\\bullet$ is $2$ or $3$ and $g$ is the explicit influence-function-type weight given in formulas (5) and (8). Under that expansion plus a central limit theorem for the integral, the standardized estimator converges to $N(0,1)$ with asymptotic variance $V_N^2 = N^{-2}\\operatorname{var}(\\sum_{i\\in s} d_i g(y_i))$, where $d_i$ are Hájek or calibration weights. The paper then provides Horvitz–Thompson and Sen–Yates–Grundy estimators for this variance. The simulation evidence is presented as supporting the claim: confidence intervals for $b_2(0.75)$ and $b_3$ tend to cover at or above the nominal rate, even though the variance estimators themselves show sizable relative bias.","pith_inferences":["If the high-level assumption holds across a wider class of designs and populations, the same linearization recipe could be carried over to other quantile-based shape measures, such as kurtosis or tail-weight indices, as long as their influence functions are available.","The bias pattern in the variance estimators suggests that a bootstrap-calibrated version of the interval, or the variance-stabilizing transformation the paper mentions, would likely deliver more accurate coverage at small sample sizes.","A quick way to stress-test the central claim is to check whether the remainder $R_N$ in the von Mises expansion is actually negligible under designs with highly unequal inclusion probabilities; simulation studies varying the spread of the $\\pi_i$ would settle this.","Since the paper leaves sufficient conditions for B1 and B2 open, bridging that gap with uniform quantile-process results would turn a heuristic argument into a theorem."],"forward_implications":["Normal confidence intervals for $b_2(r)$ and $b_3$ become available for fixed-size survey designs, with the Sen–Yates–Grundy variance estimator as the default.","The calibration cdf estimator offers little improvement over the Hájek estimator for skewness in the simulations, in contrast with the large gains it delivers for the mean; auxiliary information does less work for these functionals.","Variance estimates inherit bias from plug-in density estimation, and the relative bias need not shrink from $n=40$ to $n=80$; users should treat interval lengths as approximate at moderate sample sizes.","For the mean, the same designs show undercoverage, while the skewness intervals overcover; inference for these skewness measures is not a carbon copy of mean inference."],"supporting_citations":[{"why":"Supplies the calibration estimator and the weighted residual variance approximation used for the calibration plug-in estimators.","marker":"Deville and Särndal (1992)"},{"why":"Defines the b3 index as a global skewness measure and gives its mean-median-absolute-deviation representation.","marker":"Groeneveld and Meeden (1984)"},{"why":"Introduces the generalized Bowley measure b2(r) estimated in this paper.","marker":"Hinkley (1975)"},{"why":"Source of the design-based density estimation method used to estimate f(nu_r) by equating normal and Woodruff confidence interval lengths.","marker":"Francisco and Fuller (1986)"},{"why":"Provides the relative bias and relative stability criteria for variance estimators and contributes to the density estimation method.","marker":"Kovar et al. (1988)"},{"why":"Contributes to the density and quantile estimation method and motivates alternative cdf estimators using auxiliary information.","marker":"Rao et al. (1990)"},{"why":"Supports the use of the Sen-Yates-Grundy variance estimator for fixed-size designs.","marker":"Vijayan (1975)"},{"why":"Cited as the kind of asymptotic result that could prove the unverified conditions B1 and B2.","marker":"Conti and Marella (2015)"},{"why":"Cited as a possible route to proving the high-level assumptions for complex sampling designs.","marker":"Han and Wellner (2021)"},{"why":"Cited as another candidate quantile-process result for proving B1 and B2 in finite populations.","marker":"Dey and Chaudhuri (2024)"}],"fun_headline_variants":["Skewness measures get normal intervals in surveys","Variance formulas for Bowley and b3 skewness","Linearization gives normal inference for skewness","Survey skewness estimates now have variance estimators","Functional delta method yields skewness confidence intervals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the von Mises expansion holds with a negligible remainder and the leading linear term obeys a central limit theorem for the discrete Hájek and calibration cdf estimators; the paper states it does not prove sufficient conditions. If that assumption fails, the normal confidence intervals are not justified.","fun_headline_variants_meta":{"raw":{"variants":["Skewness measures get normal intervals in surveys","Variance formulas for Bowley and b3 skewness","Linearization gives normal inference for skewness","Survey skewness estimates now have variance estimators","Functional delta method yields skewness confidence intervals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000383,"raw_usage":{"total_tokens":1985,"prompt_tokens":858,"completion_tokens":1127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":1057}},"tokens_in":474,"tokens_out":1127,"duration_ms":8999,"temperature":1.0,"reasoning_tokens":1057,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:04:34.602667+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a finite population with a known skewness value and a design with strongly unequal inclusion probabilities, then compute $\\widehat{b}$ and the variance estimate over many samples and check the empirical distribution of $(\\widehat{b} - b)/\\widehat{V}$. If the 95% interval covers in well under 90% of samples, the central claim is false for that setting. A more direct check is to compute the remainder $R_N$ from the von Mises expansion for increasing $N$ and see whether it vanishes relative to the integral term.","supporting_citations":[{"cited_title":"and Särndal, C.-E","cited_arxiv_id":null,"evidence_quote":"Supplies the calibration estimator and the weighted residual variance approximation used for the calibration plug-in estimators."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the b3 index as a global skewness measure and gives its mean-median-absolute-deviation representation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the generalized Bowley measure b2(r) estimated in this paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the design-based density estimation method used to estimate f(nu_r) by equating normal and Woodruff confidence interval lengths."},{"cited_title":"G., Rao, J","cited_arxiv_id":null,"evidence_quote":"Provides the relative bias and relative stability criteria for variance estimators and contributes to the density estimation method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes to the density and quantile estimation method and motivates alternative cdf estimators using auxiliary information."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the use of the Sen-Yates-Grundy variance estimator for fixed-size designs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the kind of asymptotic result that could prove the unverified conditions B1 and B2."},{"cited_title":"and Wellner, J","cited_arxiv_id":null,"evidence_quote":"Cited as a possible route to proving the high-level assumptions for complex sampling designs."},{"cited_title":"and Chaudhuri, P","cited_arxiv_id":null,"evidence_quote":"Cited as another candidate quantile-process result for proving B1 and B2 in finite populations."}],"review_version":1}