{"id":"d03533d2-33d8-43dc-8921-2a74d21077b5","arxiv_id":"2412.12405","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Generalized entropy calibration, with a two-step debiasing constraint, yields doubly robust estimates of population totals from voluntary survey samples.","lead":"This paper proposes a unified calibration-weighting framework for survey data collected through voluntary participation, where inclusion probabilities are unknown. It shows how to combine propensity-score and outcome-regression models to reduce selection bias, using a two-step calibration estimator and an R package.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Two-step GEC efficiency gain is unproven: Step 2 calibrates on z=(x2, g2(ω1)), not x2 alone, so the stated reason is inaccurate and no variance comparison supports 'greater efficiency'.","rationale":"The reader's MAR concern is real but applies to the whole framework and cannot be settled by an internal check. The efficiency-of-two-step claim is a concrete gap: the estimator's stated reason is contradicted by its own calibration constraints. This is not a claim that the result is false; it is a missing proof. The paper does provide a coherent proof of Theorem 1, a real data application, and an R package, which are independent supporting evidence, but they do not address the efficiency gap. I would keep the paper CONDITIONAL: the authors should either add a variance comparison or soften the statement. Agreement with reader is partial because I did not select MAR as the primary attack; I agree it is the fundamental behavioral assumption, but the efficiency gap is more easily verified. There are also secondary proof gaps noted by the reader, such as the borrowed λ* result and unspecified regularity conditions, which reinforce the conditional verdict.","tokens_in":15922,"tokens_out":12636,"duration_ms":118886,"concrete_test":"Implement the Section 5 setting with a correctly specified reduced outcome model y_i=x_{2i}^Tβ+e_i and a logistic PS model π(x_{1i}^Tφ0), where x1 includes at least one covariate not in x2. Compute the two-step GEC estimator and two one-step GEC estimators calibrated on (i) x2 only and (ii) the full x vector over 5000 Monte Carlo repetitions. If the Monte Carlo variance of the two-step estimator is not strictly smaller than that of the x2-only one-step estimator, the claimed 'greater efficiency' of two-step calibration is refuted; if it is smaller, the current proof still needs a variance comparison theorem to justify the mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 defines the two-step estimator with Step 2 constraints (4.7)-(4.8): calibration on x2i and on g2(ω̂1i). The final weight is ω̂2i=ρ_2^{(1)}(ẑ_i^T λ̂) with ẑ_i=(x2i, ĝ2i). The paper then says 'Because only the true covariates x2 are directly used in the calibration process, the two-step GEC estimator achieves greater efficiency than the original GEC estimator when the reduced outcome model (4.2) is correctly specified.' The premise is inaccurate: g2(ω1) is itself used as a calibration covariate, so the second-step weights are not a function of x2 alone. The displayed variance under the outcome model, (1/N)Σ_{i∈S} ω_{2i}^{*2}σ^2 − σ^2, depends on the full z-calibration; no theorem or calculation shows this is no larger than the corresponding one-step variance. This matters because the paper's own Section 3 discussion after (3.9) warns that adding irrelevant calibration variables inflates exactly this variance term, and g2 is typically not in the span of x2 (it is a nonlinear function of the first-step propensity score). The efficiency claim is therefore asserted, not derived, and cannot be checked from the supplied linearization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified calibration-weighting framework for voluntary (non-probability) surveys under a missing-at-random assumption. It introduces generalized entropy calibration (GEC) as a class of weight functions, derives a dual regression-imputation representation, and develops a two-step estimator that combines a propensity-score calibration step with an outcome-regression calibration step. The central claims are that the GEC estimator is doubly robust (consistent if either the outcome regression model or the propensity score model is correctly specified), that the two-step version improves efficiency when the reduced outcome model holds, and that variance estimation is doubly robust. The results are supported by a small simulation using a Korean health database and a real-data integration example.","tokens_in":16219,"tokens_out":10296,"duration_ms":87451,"significance":"If correct, the framework is a useful unification: it puts several existing calibration and propensity-score methods under one convex-duality umbrella, provides a clear link to regression estimation, and offers an implementable two-step procedure for survey nonresponse adjustment. The paper also ships an R package (GECal) and gives a real-data illustration, which are practical strengths. The dual relationship and the debiasing constraint idea are valuable. However, the significance is tempered by the fact that the main theorems rely on unstated regularity conditions, the identification step for the propensity-model branch is delegated to an external result, the claimed efficiency gain of the two-step estimator is not proven, and the variance derivation under the outcome model appears to contain an unjustified simplification.","major_comments":[{"comment":"The claim that the two-step GEC estimator achieves greater efficiency because 'only the true covariates x2 are directly used in the calibration process' is inaccurate. The second-step constraints in (4.7)-(4.8) calibrate on z=(x2, g2(ω̂1)), not on x2 alone. The displayed variance for the two-step estimator depends on the full z-calibration, and no theorem or calculation shows that this variance is no larger than the corresponding one-step variance. This matters because the paper's own discussion after (3.9) warns that irrelevant calibration variables inflate the variance, and g2(ω̂1) is typically not in the span of x2. The efficiency claim is therefore asserted, not derived.","section":"Section 4, Eq. (4.10) and following paragraph"},{"comment":"Both theorems are stated under 'some regularity conditions' that are never enumerated. The asymptotic linearizations (3.5) and (4.12) with o_p(n^{-1/2}N) rates require conditions on the entropy function G, the covariate distributions, the moments of the study variable, and the existence and uniqueness of the probability limits λ* and γ*. Without these conditions, the double-robustness and local-efficiency claims are not verifiable, and the theory is not self-contained.","section":"Theorems 1 and 2"},{"comment":"The key identification λ*=(0^T,1)^T under the propensity-score model is delegated to 'the same argument for proving Theorem 1 in Lesage et al. (2019)' without stating the conditions or reproducing the argument. The subsequent two-stage Taylor expansion is only sketched; in particular, the term involving ẑ_i^T γ*_2 in (A.10) and the claim that the second term has zero expectation require careful handling of the randomness in the first-step estimate. As written, the propensity-model branch of the double-robustness result is not established within this manuscript.","section":"Appendix, proof of Theorem 2"},{"comment":"The variance expression under the outcome-regression model is not justified. The text replaces E[(δ_i ω_i^* −1)^2 σ^2] with (1/N)Σ_{i∈S}(ω_i^{*2} − ω_i^*)σ^2, attributing the last step to the calibration constraint Σ_{i∈S} ω̂_i=N. However, E[(δ_i ω_i^* −1)^2]=E[π_i ω_i^{*2} −2π_i ω_i^* +1], and the calibration constraint does not remove the dependence on π_i. The simplification holds only if π_i=1/ω_i^*, i.e., under the propensity-score model, not under the outcome-regression model alone. Consequently, the doubly robust variance estimator in (3.11) is not supported by the stated argument, and the simulation coverage rates for the OR-model scenarios rely on an unproven variance claim.","section":"Section 3, Eq. (3.9)"}],"minor_comments":[{"comment":"There are typographical errors: 'calibraton' should be 'calibration', and 'similiar' should be 'similar'.","section":"Section 3, text after Eq. (3.4) and after Eq. (3.16)"},{"comment":"The notation for z_i is inconsistent: the display after (4.10) uses z_i^⊤=(x_{2i}^⊤, g_{2i}), while earlier ẑ_i is defined with the estimated ĝ_{2i}; please clarify the distinction between z_i and its estimated counterpart.","section":"Section 4, Eqs. (4.10)-(4.12)"},{"comment":"The simulation does not include a scenario that isolates the claimed efficiency advantage of the two-step estimator over the one-step estimator under the same propensity-score model; the comparison of PS2 and PS3 changes both the propensity model and the number of steps, so it does not directly test the efficiency claim.","section":"Section 5, Table 3"},{"comment":"The standard errors of the GEC estimates are computed by treating the calibrated weights as fixed within the survey package; the uncertainty from estimating the calibration parameter λ is not accounted for, which may understate the standard errors.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper builds directly on the authors' prior work (Kwon et al., 2024) and on instrumental-calibration arguments of Lesage et al. (2019), and the incremental novelty should be clarified. The propensity-score model in (3.8) is defined as the inverse of the calibration weighting function, so the PS branch of double robustness is close to being a consequence of the estimator's normalization; the proof needs to show more than the fact that the calibration equations are satisfied. The claimed local efficiency is also only heuristic, with no formal efficiency theorem stated. These points, together with the unproven two-step efficiency gain and the questionable variance derivation, suggest that the manuscript requires substantial revision before the central claims are fully supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the Kwon-Kim-Qiu paper on generalized entropy calibration for voluntary samples. The core idea is solid: they unify a broad class of calibration weights through convex conjugates, give a calibration generating function that matches any monotone inverse-link propensity model, and add a two-step version where the first step builds propensity weights and the second step calibrates on outcome-model covariates plus a debiasing constraint. That gives a doubly robust estimator under OR or PS models, with a plausible linearization. The R package GECal is a plus; the simulations show the two-step estimator tracks the oracle performance well and the real-data example is a sensible illustration.\n\nThe main soft spot is the efficiency claim. The paper says the two-step estimator beats the one-step version when the reduced outcome model is correct because 'only the true covariates x2 are directly used in the calibration process.' But the second step calibrates on z = (x2, g2(ω1)), and the debiasing constraint itself enters the weight. The variance expression depends on ω2* which is a function of the whole z. Nothing in the paper shows this variance is no larger than the one-step version's, and their own Section 3 warns that extraneous calibration variables inflate exactly this term. So that claim is asserted, not derived. It needs a proof or a more careful statement. As a minor point, the 'regularity conditions' for Theorems 1 and 2 are unspecified, and Theorem 2's λ* = (0,1) argument is delegated to Lesage et al. (2019); a self-contained derivation would help.\n\nI also think the PS model definition in (3.8) deserves one sentence of clarification. It is defined as the inverse of the calibration weight, so the PS branch of double robustness is literally baked into the weight construction. That is not a fatal flaw—it's how calibration works—but it should be stated so readers don't mistake it for an independently specified model.\n\nBottom line: the paper deserves a serious referee. The unification and the two-step construction are genuine contributions, and the deficiencies are fixable. If I were handling it, I'd ask for a variance comparison (analytic or simulation) for the efficiency claim, a spelled-out regularity conditions statement, and a clearer discussion of the PS-model circularity.\n\nWould I cite it? Probably yes if I were working on non-probability sample calibration. Reading group: maybe.","headline":"A useful unified calibration framework with a credible double-robustness story, but the two-step efficiency claim rests on an unstated premise and needs a proper derivation or a softer claim.","tokens_in":16703,"tokens_out":3174,"would_cite":true,"duration_ms":28163,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that generalized entropy calibration, applied in two steps with a propensity-score weight and a debiasing outcome-regression constraint, produces a doubly robust, locally efficient estimator for voluntary survey data.","keywords":["generalized entropy calibration","calibration weighting","voluntary survey data","nonprobability samples","doubly robust estimation","propensity score","outcome regression","two-step calibration"],"falsifier":"Generate a population in which the outcome regression and the propensity-score model are both correctly specified as in (2.3) and (3.8), draw many voluntary samples, and compute the two-step GEC estimate and the variance estimator in (3.11); if the bias does not shrink at the $n^{-1/2}$ rate or the 95% intervals do not stay near nominal coverage under either correct model, the double-robustness claim is not supported.","tokens_in":15733,"feed_emoji":"📊","tokens_out":11113,"duration_ms":96748,"temperature":0.7,"pith_summary":"This paper proposes generalized entropy calibration (GEC) as a unified way to reweight voluntary survey samples when the sampling mechanism is ignorable. It shows that any GEC estimator has a dual representation as a regression estimator, and that a two-step version—first calibrating on a working propensity-score model, then calibrating on outcome-regression covariates with a debiasing constraint—is consistent if either the outcome model or the propensity-score model is correctly specified. The practical payoff is that analysts can remove selection bias using only population totals of auxiliary variables, without needing individual-level covariate data for the whole population. The paper also derives asymptotic linearizations and a doubly robust variance estimator, and demonstrates the method on simulations and a data-integration application.","feed_headline":"Fix voluntary-survey bias with one of two working models","feed_subtitle":"Generalized entropy calibration stays unbiased if either the outcome or propensity model is correct, using only population totals.","key_machinery":"The engine of the method is the pair $(G,\\rho)$: $G$ is a strictly convex differentiable entropy function, and $\\rho$ is its convex conjugate, whose derivative $\\rho^{(1)}$ maps the Lagrange multiplier $\\lambda$ to the calibration weights. The dual objective identifies the regression model implicit in any choice of $G$, which is what allows model selection for calibration. In the two-step version, the first step uses a calibration generating function built from the working propensity-score link via $\\rho^{(1)}(\\nu)=1/\\pi(\\nu)$, and the second step calibrates on the outcome-regression covariates plus the debiasing constraint $\\sum_{i\\in S} \\omega_i g_2(\\hat\\omega_{1i}) = \\sum_{i=1}^N g_2(\\hat\\omega_{1i})$, which is the constraint that makes the final estimator consistent under the propensity-score model.","core_discovery":"The central claim is that the generalized entropy calibration estimator, including the two-step version, is doubly robust and locally efficient. For a strictly convex entropy $G$, the weights minimize $\\sum_{i\\in S} G(\\omega_i)$ subject to the calibration constraint $\\sum_{i\\in S} \\omega_i x_i = \\sum_{i=1}^N x_i$; the dual optimization through the convex conjugate $\\rho$ gives weights $\\hat\\omega_i = \\rho^{(1)}(x_i^\\top \\hat\\lambda)$ and a linearized estimator that is a weighted regression prediction plus a debiasing residual term. The paper's Theorem 1 states this linearization without needing any model to be correct. Consistency then follows if either the linear outcome regression model or the propensity-score model is correctly specified, and under the propensity-score model the estimator reaches the Godambe-Joshi lower bound. The two-step estimator first uses an entropy whose conjugate derivative is the inverse of a chosen propensity-score link, then adds a debiasing calibration constraint; when the reduced outcome model is correct, the two-step estimator is more efficient than the one-step version.","pith_inferences":["A testable extension would map the bias-variance trade-off between one-step and two-step calibration as the outcome model is deliberately misspecified; the paper compares oracle, constant, and estimated propensity scores, but not a full misspecification grid.","The empirical similarity across entropies suggests the entropy choice matters mainly through weight boundedness and convergence, so a practitioner could default to a simple entropy and treat the weight bound as the real tuning parameter.","Because the identifying assumption is missingness at random, the method's practical value is bounded by how well the auxiliary variables capture the drivers of volunteering; editors might pair GEC with sensitivity analyses for unmeasured selection.","The dual relationship implies that regularization techniques for high-dimensional regression could be imported into calibration weighting, an extension the paper flags but does not develop."],"forward_implications":["A survey analyst who trusts either a propensity model for participation or a linear model for the study variable can use the same two-step calibration and obtain consistent estimates.","The variance estimator in (3.11) stays valid when either of the two working models is correct, so confidence intervals inherit the double robustness.","When the reduced outcome regression model is right, the two-step estimator has smaller asymptotic variance than one-step calibration, so adding the outcome model buys efficiency rather than just bias protection.","Because the method needs only population totals of auxiliary variables, it applies to privacy-sensitive contexts in which unit-level population covariates are unavailable.","The construction extends directly to multiply robust calibration by adding further working models, as the paper notes in its concluding remarks."],"supporting_citations":[{"why":"Establishes the Lagrange-multiplier calibration construction whose weights the GEC method generalizes.","marker":"Deville and Särndal (1992)"},{"why":"Defines the missingness-at-random condition (2.1) that is the method's identifying assumption.","marker":"Rubin (1976)"},{"why":"Supplies the debiased calibration weighting and variance formula that the two-step estimator and (3.11) build on.","marker":"Kwon et al. (2024)"},{"why":"Provides the calibrated maximum likelihood approach that appears as a special case of the calibration generating function with logistic propensity scores.","marker":"Tan (2020)"},{"why":"Gives the information-projection perspective used in Remark 1 and Lemma 3 for conditional independence under the reduced model.","marker":"Wang and Kim (2024)"},{"why":"Frames the doubly robust estimation concept that the GEC estimator's consistency under either model is compared against.","marker":"Bang and Robins (2005)"},{"why":"Supplies the argument used in the proof of Theorem 2 for the limiting value of the second-step Lagrange multiplier.","marker":"Lesage et al. (2019)"},{"why":"Shows how to extend to multiply robust estimation, which the paper cites for the multivariate extension of GEC.","marker":"Han (2014)"}],"fun_headline_variants":["Doubly robust calibration for voluntary surveys","One calibration, two models: robust survey weighting","Two-step calibration trims survey bias efficiently","Generalized entropy calibration: robust to model misspecification","Voluntary survey? Use this doubly robust estimator"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is missingness at random: after conditioning on the observed covariates, the decision to volunteer must carry no extra information about the study value, because otherwise no calibration weight built from those covariates can remove the selection bias.","fun_headline_variants_meta":{"raw":{"variants":["Doubly robust calibration for voluntary surveys","One calibration, two models: robust survey weighting","Two-step calibration trims survey bias efficiently","Generalized entropy calibration: robust to model misspecification","Voluntary survey? Use this doubly robust estimator"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1302,"prompt_tokens":884,"completion_tokens":418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":348}},"tokens_in":500,"tokens_out":418,"duration_ms":4253,"temperature":1.0,"reasoning_tokens":348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:06:58.833810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a population in which the outcome regression and the propensity-score model are both correctly specified as in (2.3) and (3.8), draw many voluntary samples, and compute the two-step GEC estimate and the variance estimator in (3.11); if the bias does not shrink at the $n^{-1/2}$ rate or the 95% intervals do not stay near nominal coverage under either correct model, the double-robustness claim is not supported.","supporting_citations":[{"cited_title":"and S¨ arndal, C.-E","cited_arxiv_id":null,"evidence_quote":"Establishes the Lagrange-multiplier calibration construction whose weights the GEC method generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the missingness-at-random condition (2.1) that is the method's identifying assumption."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the calibrated maximum likelihood approach that appears as a special case of the calibration generating function with logistic propensity scores."},{"cited_title":"and Kim, J","cited_arxiv_id":null,"evidence_quote":"Gives the information-projection perspective used in Remark 1 and Lemma 3 for conditional independence under the reduced model."},{"cited_title":"and Robins, J","cited_arxiv_id":null,"evidence_quote":"Frames the doubly robust estimation concept that the GEC estimator's consistency under either model is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the argument used in the proof of Theorem 2 for the limiting value of the second-step Lagrange multiplier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows how to extend to multiply robust estimation, which the paper cites for the multivariate extension of GEC."}],"review_version":1}