{"id":"100159e0-edb0-45ac-9383-ff10047b37f5","arxiv_id":"2606.31381","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Model-assisted calibration integrates multiple probability surveys to boost regression efficiency while preserving design-based finite-population inference.","lead":"The paper proposes model-assisted calibration methods to integrate data from multiple probability-based surveys and improve regression efficiency while handling complex sampling designs. A smart generalist might read it to see how existing public survey data can be combined for more precise population estimates without new data collection.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest_assumption (common target population) is stated by the paper as naturally satisfied and is not load-bearing for the internal validity of the consistency or variance results. Full-text derivations would be needed to confirm technical details, but nothing in the abstract or claim description indicates a soft spot that would alter the UNVERDICTED status.","tokens_in":1696,"tokens_out":293,"duration_ms":26452,"concrete_test":"Re-derive the asymptotic design consistency of the proposed estimator (presumably Theorem 1 or equivalent) under the case where the external survey supplies only summary statistics; confirm that the resulting estimating equation still reduces to the usual calibration constraint when the two sampling frames coincide.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that model-assisted calibration estimators integrating multiple probability surveys are design-consistent for finite-population regression parameters, remain valid without correct outcome-model specification, and yield efficiency gains while properly accounting for both surveys' complex designs (via Taylor linearization). This is a direct extension of standard model-assisted survey estimation (e.g., generalized regression estimators) to the multi-survey setting; the common-target-population assumption is explicitly justified by the probability-sampling property and does not introduce an internal inconsistency. No derivation gap, hidden model dependence, or unaccounted design feature is apparent from the stated framework.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes model-assisted calibration methods to integrate multiple probability-based survey samples for improving the efficiency of finite-population regression analyses. The framework accommodates both individual-level data and summary statistics from external surveys, establishes design consistency of the estimators, develops Taylor linearization variance estimators that account for the complex designs of both surveys, and does not require correct specification of the outcome model. Efficiency gains are demonstrated through simulation studies and an application integrating NHANES and NHIS data.","tokens_in":1820,"tokens_out":543,"duration_ms":22994,"significance":"If the central results hold, the work extends standard model-assisted survey estimation (e.g., generalized regression estimators) to a multi-survey integration setting while preserving design-based finite-population inference. This is particularly relevant for public health and social science applications that routinely combine representative probability surveys. The explicit handling of summary statistics only, the design-consistency proofs, and the Taylor linearization variances that incorporate both sampling designs are notable strengths; the simulations and real-data example provide concrete evidence of efficiency gains without sacrificing validity.","major_comments":[{"comment":"§4.2, Theorem 1: the design-consistency argument appears to rely on the common target population being exactly the same for both surveys; while the abstract states this is naturally satisfied, the proof sketch should explicitly address the case where the two sampling frames have minor but non-negligible overlap differences, as this could affect the bias term in the linearization.","section":"§4.2, Theorem 1"},{"comment":"§5.3, Eq. (18): the Taylor linearization variance estimator is stated to account for both designs, but the cross-term arising from the calibration weights estimated from the external survey is not shown explicitly; if this term is omitted, the reported variances may be understated when the external sample size is moderate.","section":"§5.3, Eq. (18)"}],"minor_comments":[{"comment":"Table 2: the column labels for the 'summary statistics only' scenario are not fully aligned with the notation introduced in §3.1; adding a footnote linking the columns to the relevant equations would improve readability.","section":"Table 2"},{"comment":"The reference list omits several recent papers on multi-frame survey calibration (e.g., works extending the generalized regression estimator to multiple frames); including 2–3 such citations in the introduction would better situate the contribution.","section":"Introduction"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive review and positive assessment of our work. The comments have prompted useful clarifications. We respond to each major comment below.","responses":[{"response":"We appreciate this observation. Theorem 1 is derived under the maintained assumption that both probability surveys target the identical finite population, which holds by design for surveys such as NHANES and NHIS. Under standard survey asymptotics, minor frame discrepancies would contribute a bias term of smaller order than the leading variance terms. To improve transparency, we will revise the proof sketch in §4.2 to state the common-population assumption explicitly and add a brief remark on the approximation that applies when frame overlap is nearly complete.","revision_made":"yes","referee_comment":"[§4.2, Theorem 1] §4.2, Theorem 1: the design-consistency argument appears to rely on the common target population being exactly the same for both surveys; while the abstract states this is naturally satisfied, the proof sketch should explicitly address the case where the two sampling frames have minor but non-negligible overlap differences, as this could affect the bias term in the linearization."},{"response":"We thank the referee for noting this presentational detail. The linearization underlying Eq. (18) is obtained from the joint influence function of the calibration estimator and therefore includes the cross-term that arises from estimating the calibration weights on the external sample. The term was suppressed in the displayed expression for brevity. In the revision we will expand Eq. (18) to display the cross-term explicitly, confirming that the variance estimator remains design-consistent for the joint sampling process.","revision_made":"yes","referee_comment":"[§5.3, Eq. (18)] §5.3, Eq. (18): the Taylor linearization variance estimator is stated to account for both designs, but the cross-term arising from the calibration weights estimated from the external survey is not shown explicitly; if this term is omitted, the reported variances may be understated when the external sample size is moderate."}],"tokens_in":1384,"tokens_out":451,"duration_ms":28336,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper extends model-assisted calibration to integrate multiple probability surveys that have complex sampling designs. The setup lets researchers use either individual-level data or just summary statistics from the external survey, produces design-consistent estimators for finite-population regression parameters, and supplies Taylor linearization variances that account for both designs. It shows efficiency gains in simulations and in an NHANES-NHIS application.\n\nWhat is new is the adaptation to the multi-survey probability setting with complex designs; most prior integration work focused on nonprobability samples and did not handle this case. The common-target-population assumption holds naturally for these surveys, and the framework avoids requiring the outcome model to be correctly specified.\n\nThe paper does the standard things well: it states design consistency, develops the variance estimators, and demonstrates practical gains. The extension follows directly from generalized regression estimation ideas applied to two samples.\n\nSoft spots are minor. The abstract gives no equations, so the exact form of the calibration weights and how summaries enter the estimator cannot be checked here, but nothing in the description suggests hidden model dependence or unaccounted design features. When only summaries are available the efficiency gain is likely smaller, and it would help to see how much is lost relative to full data. The variance estimator when summaries are used also needs clear implementation details.\n\nThis is for survey statisticians and applied researchers who combine national probability surveys. Readers working on finite-population inference or data integration will find the methods and the empirical results useful. It deserves a serious referee because the claims rest on established survey theory, the problem is relevant, and the evidence from simulations plus the real-data example is concrete.","headline":"Extends model-assisted calibration to multi-survey integration with complex designs, allowing summary stats or full data while keeping design consistency and valid inference without correct outcome model.","tokens_in":2262,"tokens_out":407,"would_cite":false,"duration_ms":27539,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Model-assisted calibration integrates multiple probability surveys to raise regression efficiency while preserving design-based validity.","keywords":["data integration","model-assisted calibration","survey sampling","regression efficiency","finite population inference","probability samples","complex sampling designs"],"falsifier":"A simulation or real-data check in which the integrated estimator exhibits large bias when the two surveys are known to target populations with different covariate distributions would falsify the design-consistency claim.","tokens_in":2602,"feed_emoji":"","tokens_out":598,"duration_ms":24329,"temperature":0.7,"pith_summary":"The paper develops calibration techniques that borrow strength across separate probability surveys, each drawn from the same target population, to produce more precise regression estimates. The methods accept either full individual records or only summary statistics from an external survey and do not require the working regression model to be correctly specified. Design consistency is maintained because the calibration respects the known sampling probabilities of each survey. Taylor linearization yields variance estimates that properly reflect the complex designs of all surveys involved. Real-data illustration with two national health surveys shows measurable precision gains without sacrificing finite-population inference.","feed_headline":"Calibration merges survey samples to sharpen regression estimates","feed_subtitle":"Method keeps design-based validity and works with either full records or summary statistics alone.","key_machinery":"Model-assisted calibration that uses known or estimated population totals from an auxiliary probability survey to adjust the regression estimator while respecting each survey's sampling design.","core_discovery":"The central claim is that model-assisted calibration estimators, constructed by adjusting regression residuals with auxiliary totals or estimates drawn from a second probability survey, are design-consistent for finite-population regression parameters and asymptotically more efficient than the single-survey estimator, even when the outcome model is misspecified.","pith_inferences":["The same calibration logic could be applied when one survey supplies only cell-level means rather than microdata, lowering data-sharing barriers.","If future surveys adopt compatible sampling frames, repeated application of the method across waves would accumulate efficiency gains over time.","The approach supplies a concrete way to test whether efficiency improvements persist when the auxiliary survey covers only a subset of the covariates used in the outcome model."],"forward_implications":["The resulting estimators remain consistent under the joint sampling design of the surveys involved.","Variance estimators obtained by Taylor linearization account for the complex sampling of both surveys and remain valid under model misspecification.","The framework covers both the case of full individual-level data and the case of only summary statistics from the external survey.","Application to NHANES and NHIS data produces regression estimates with visibly smaller standard errors than either survey alone."],"fun_headline_variants":["Model-assisted calibration integrates probability surveys for regression","Calibration merges survey samples to improve regression estimates","Integrating multiple probability surveys via model-assisted calibration","Model-assisted calibration of regression using dual probability surveys"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The separate probability surveys are each designed to represent the identical target population.","fun_headline_variants_meta":{"raw":{"variants":["Model-assisted calibration integrates probability surveys for regression","Calibration merges survey samples to improve regression estimates","Integrating multiple probability surveys via model-assisted calibration","Model-assisted calibration of regression using dual probability surveys"]},"model":"grok-4.3","cost_usd":0.01039,"raw_usage":{"total_tokens":4576,"prompt_tokens":625,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":103899500,"prompt_tokens_details":{"text_tokens":625,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3896,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":625,"tokens_out":55,"duration_ms":46549,"temperature":1.0,"reasoning_tokens":3896,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T04:36:45.029869+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A simulation or real-data check in which the integrated estimator exhibits large bias when the two surveys are known to target populations with different covariate distributions would falsify the design-consistency claim.","supporting_citations":[],"review_version":1}