{"id":"4f86d317-017b-4726-ba28-8a9d1f7dc72b","arxiv_id":"2501.04185","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new pseudo-likelihood estimator, Overlap-Excluded, uses survey units not present in the big dataset to estimate selection propensities and produces more efficient inverse-probability-weighted estimates in simulations.","lead":"This paper proposes a new way to estimate how likely each person or business is to appear in a big, non-random dataset, by combining the big data with a smaller, carefully drawn survey. The method, called Overlap-Excluded propensity estimation, aims to make population estimates from big data less biased and more precise.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 6 admits the efficiency proof is not included, and the promised Poisson-sampling proof would not match the PPS reference design used in the Section 5 simulation; the central efficiency claim currently rests on one well-specified DGP.","rationale":"The paper's central claim has two parts: unbiasedness under a correctly specified model, and efficiency relative to alternatives. Unbiasedness is plausible because the OE score in Eq. (13) is an unbiased estimating equation under correct specification and independence of A and B: E[S_OE(theta0)] = sum_{U} pi(1-pi)x - sum_{U} pi(1-pi)x = 0. The problem is the efficiency part. Section 6 states that a theorem is available but 'not included in this paper', and it specifies Poisson sampling for the reference sample, while the simulation in Section 5 uses PPS sampling for A. Thus the only direct evidence for the headline efficiency ranking is a narrow simulation with a single correctly specified DGP, one population, one reference sample size, and perfect linkage. The reader's weakest assumption (accurate linkage) is an important scope condition that applies mainly to implementation; the more immediate gap is that the paper itself concedes the theoretical support for its main efficiency claim is absent. The proposed check, deriving the asymptotic variance under the actual PPS design or verifying whether the Poisson-based proof extends, would settle whether the efficiency claim is genuine or an artifact of the single simulated setting. If the proof cannot be produced, the paper should be accepted only with the efficiency claim downgraded to 'observed in simulation for the settings considered.'","tokens_in":18,"tokens_out":6765,"duration_ms":129913,"concrete_test":"Independently derive the asymptotic variance of the OE IPW estimator under the PPS reference design actually used in Section 5 (randomized systematic PPS with size variable c + x3), and compare it with the corresponding variances for the CLW, KW, and WVL estimators. If the OE variance is not uniformly no larger than that of CLW and KW, or if the derivation cannot be completed for the PPS design, then the Section 6 efficiency assertion is unsupported for the simulated design; at a minimum, the claim should be restricted to the Poisson-sampling case and the simulation rerun under Poisson A.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 states: 'While not included in this paper, it can be shown theoretically (confirming the simulation results in this paper) that when the reference sample is taken by Poisson sampling the OE IPW estimator is at least as efficient as the Hájek-like estimator from Chen et al. (2020) and the Kim & Wang (2019) data integration estimator.' Two problems follow. First, no proof or reference is actually provided, so the paper's headline claim that OE is more efficient than existing methods, particularly for large B, is supported only by the simulation in Section 5. Second, the promised proof is expressly for a Poisson-sampled reference sample, whereas Section 5 says the reference sample A was selected by 'randomised systematic probability proportional to size sampling (PPS)'. Thus the theoretical statement, even if true, would not verify the simulation results actually reported. The simulation itself covers one logistic DGP, one reference sample size (nA = 5, 000), one population, and perfect linkage. The efficiency ranking (OE best at NB = 50, 000 and 140, 000, WVL best at NB = 2, 000) is therefore not established outside that setting. The plug-in variance estimator of Eqs. (17)-(19), with alternative (20), is also evaluated only under the same correctly specified model; its behavior under misspecification or linkage errors is untested. The conditional verdict is warranted, but the missing proof and the design mismatch are the key gaps.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an Overlap-Excluded (OE) propensity score estimator for correcting selection bias when estimating a finite-population mean from a non-probability sample B with the help of a probability reference sample A. The method assumes that membership of units in A in B can be identified exactly. The OE pseudo-likelihood in Eq. (11) uses B for the positive term and A\\B, weighted by the survey design weights, for the negative term. The paper derives the corresponding score and Newton-Raphson updating equations, defines an inverse probability weighted estimator (16), and presents plug-in variance estimators in Eqs. (17)-(20). The method is compared with the CLW, KW, WVL, and VD estimators in a simulation study with N=200,000, nA=5,000, and NB equal to 2,000, 50,000, or 140,000. The reported simulation results show OE with low bias and the lowest RRMSE for the two larger values of NB, while WVL is best for the smallest NB. The paper concludes that when the propensity model is correctly specified, the OE-based IPW estimator is unbiased and efficient relative to the alternatives, particularly for large B.","tokens_in":8371,"tokens_out":3969,"duration_ms":41973,"significance":"If the claimed efficiency properties were rigorously established, the OE estimator would be a practically useful addition to the toolkit for integrating big non-probability data with probability samples, especially in settings where common identifiers or high-quality record linkage make the overlap between A and B known. The paper has clear strengths: the setup is clean, the simulation design is described in sufficient detail to be reproducible, and the comparison includes four established alternative estimators. The Monte Carlo results are plausible and suggest that using A\\B rather than all of A in the pseudo-likelihood can improve efficiency. However, the paper's central theoretical claims are currently only asserted, not proved, and the simulation evidence is limited to a single correctly specified logistic model, one reference-sample size, one population, and perfect linkage. Because the manuscript's headline message is an efficiency ranking, the missing proof and the limited evidence are load-bearing weaknesses. The method itself is not circular or incoherent; the gaps are in the supporting theory and the breadth of the empirical evaluation.","major_comments":[{"comment":"The concluding efficiency claim rests on an unproved assertion. The text states that 'while not included in this paper, it can be shown theoretically (confirming the simulation results in this paper) that when the reference sample is taken by Poisson sampling the OE IPW estimator is at least as efficient as' the CLW and KW estimators. No proof or reference is supplied. Since the abstract and Section 6 present efficiency relative to alternatives as a main message, this is a central claim rather than a side remark. Moreover, the promised theorem is stated for a Poisson-sampled reference sample, whereas the Section 5 simulation states that the reference sample A is drawn by 'randomised systematic probability proportional to size sampling (PPS)'. A Poisson-sampling theorem would therefore not confirm the simulation results actually reported. The authors should either provide a proof that matches the simulation design or explicitly label the efficiency claim as a conjecture supported only by the simulation.","section":"Section 6"},{"comment":"The simulation evidence is too narrow to support the general efficiency claim in Section 6. Table 1 covers a single data-generating process, a single population (N=200,000), a single reference sample size (nA=5,000), one logistic propensity model, and perfect linkage between A and B. The ranking in Table 1 (OE best for NB=50,000 and 140,000, WVL best for NB=2,000) may depend on these choices. In particular, the paper does not examine misspecification of the propensity model, varying nA, varying overlap rates, or linkage errors. At minimum, the conclusion should be restricted to the design actually simulated, or the simulation should be extended to settings that probe the claimed conditions. Without this, the sentence in Section 6 that the estimates are efficient 'relative to the other approaches we have considered' overstates what the evidence shows.","section":"Section 5"},{"comment":"The alternative plug-in variance estimator in Eq. (20) is introduced without derivation and is not evaluated in the simulation. Eq. (20) is intended to estimate the component in Eq. (19) without requiring design weights for units in B, which would be important in practice because those weights may be unavailable. The paper gives no argument for its unbiasedness or consistency, and Table 2 does not state whether the coverage probabilities were computed using Eqs. (17)-(19) or Eq. (20). The authors should either derive Eq. (20), provide a reference, or study it in the simulation; as written, the paper contains an untested variance estimator that is presented as an alternative to a component of the main variance formula.","section":"Section 4.2"},{"comment":"The unbiasedness claim in Section 6 ('When the model for π_i^B is correctly specified, the IPW estimator formed from the OE propensity scores is unbiased') is not established in the paper. Section 4 derives the pseudo-likelihood and score equation heuristically, but no formal consistency or unbiasedness theorem is given for the maximum pseudo-likelihood estimator θ-hat from Eq. (13), nor for the resulting IPW estimator (16). The simulation results at NB=2,000 show small biases, but a theoretical statement should be proved or explicitly labelled as a conjecture. The authors should add a theorem with regularity conditions, or revise the wording to avoid making an unproved theoretical assertion.","section":"Section 4 and Section 6"}],"minor_comments":[{"comment":"The title in the manuscript text has an extra space in 'D ata Integration'; this should be corrected to 'Data Integration'.","section":"Title"},{"comment":"The displayed equations for the plug-in variance estimator are split by line numbers in a way that makes them hard to read; the plus signs appear on separate lines from the terms they connect. Please reformat the display as a single equation or a clearly aligned multi-line equation.","section":"Section 4.2, Eqs. (17)-(19)"},{"comment":"The caption of Table 2 should state explicitly which of the two plug-in variance estimators, Eqs. (17)-(19) or Eq. (20), was used to construct the confidence intervals, since the paper presents both.","section":"Section 5.1, Table 2"},{"comment":"The text reports that WVL has the lowest RRMSE for NB=2,000, but does not discuss why the ranking reverses as NB grows; a brief explanation or at least an acknowledgment would help the reader interpret the headline claim.","section":"Section 5.1, Table 1"},{"comment":"The phrase 'H´ajek-like estimator' contains a typographical issue with the diacritic; the name should be rendered as 'Hájek-like'.","section":"Section 6"},{"comment":"Assumptions A1-A3 are stated but are not used in any formal derivation; the paper should either state that they are assumptions for the theoretical properties, or connect them to the estimating equation (13) and the variance estimators.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a plausible estimator and a clean simulation study, but the central efficiency claim is currently supported only by a single simulation and by an unproved assertion in Section 6. The promised Poisson-sampling theorem would not confirm the PPS-based simulation, so the authors need either to supply a matching proof or to substantially temper the claims. This is fixable within the paper's scope, so I recommend major revision rather than rejection. In revision, please also clarify which variance estimator was used in Table 2 and add at least a derivation or simulation evaluation for Eq. (20)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful small step, not a breakthrough. The idea is to estimate propensities into a non-probability big data set using a reference probability sample, but to drop the reference sample units that also appear in the big data from the non-inclusion term. That twist (A\\B instead of all A, and modeling pi directly rather than the WVL p) is real and clearly explained. The simulation is clean and the results are sensible: OE has the lowest RRMSE for large B; WVL wins at NB=2,000. Coverage of the plug-in variance is close to nominal. So the paper has value as a careful comparison and a plausible alternative estimator.\n\nThe soft spots are the ones the stress-test flags. Section 6 promises a theoretical proof of the efficiency claim but doesn't give it or a reference. That in itself is a completeness gap, not a fatal flaw, but the bigger mismatch is that the promised result is for Poisson sampling of the reference sample, while the simulation uses PPS. So even if the theorem holds, it would not verify the headline simulation. The central efficiency claim for large B rests on one well-specified DGP, one population, one nA, perfect linkage, and no misspecification. The alternative variance estimator in Eq (20) is derived without proof and not tested in the simulation, so we don't know if it preserves coverage. These are fixable with more simulation and an actual proof or a clear statement that efficiency is empirical.\n\nI was a bit worried about circularity, but there's none: the estimator is fitted to a pseudo-likelihood and evaluated on simulated external baselines. The authors are honest about the linkage assumption, and they compare against the right literature.\n\nBottom line: this deserves a serious referee. It's not ready as is — the missing proof and the design mismatch have to be addressed, and the simulation should include at least one misspecified model or linkage-error scenario. But the core idea is sound and likely useful for official statistics. I'd send it to review, and if I were working on data integration methods I'd probably cite the OE estimator even in its current form.","headline":"A clean, incremental variant of propensity weighting for linked non-probability data; the simulation is promising but the central efficiency theorem is missing and even its stated setting doesn't match the simulation design.","tokens_in":8893,"tokens_out":2803,"would_cite":true,"duration_ms":28249,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes the Overlap-Excluded propensity score estimator, and claims that with a correctly specified selection model the resulting inverse-probability-weighted estimates are unbiased and more efficient than four standard…","keywords":["big data","selection bias","propensity score estimation","inverse probability weighting","data integration","non-probability samples","reference sample","overlap-excluded estimator"],"falsifier":"In the paper's simulation design, deliberately misclassify 5% of the overlap indicators in the reference sample and compare the OE IPW estimates with those from CLW under the same corrupted linkage; if OE's bias exceeds CLW's, the accurate-linkage assumption is the limiting premise.","tokens_in":7903,"feed_emoji":"📊","tokens_out":10105,"duration_ms":82138,"temperature":0.7,"pith_summary":"Big-data sources such as administrative and sensor records cover the population unevenly, so using them directly gives biased estimates of means and totals. The paper's route to correction is to integrate the big dataset with a small probability reference sample, estimate each big-data unit's propensity to be included, then weight by the inverse of that propensity. The proposed Overlap-Excluded (OE) estimator builds the propensity model from two pieces: every unit in the big dataset contributes its own inclusion log-likelihood, and the reference units that are not also in the big dataset contribute the non-inclusion log-likelihood, weighted by their design weights. The paper claims that, under a correctly specified selection model, the OE inverse-probability-weighted estimator is unbiased and, in simulations with large big-data samples, has lower relative root mean squared error than four established alternatives while keeping coverage near nominal levels. If that holds, statistical agencies can get more precise official statistics from large non-probability datasets without collecting new probability samples.","feed_headline":"Overlap-excluded propensities sharpen big-data estimates","feed_subtitle":"Excluding overlap with a reference sample makes big-data estimates unbiased and gives the lowest error in large-data simulations.","key_machinery":"The central object is the Overlap-Excluded (OE) pseudo-likelihood, defined as $l_{OE}(\\theta) = \\sum_{i \\in B} \\log \\pi^B_i + \\sum_{i \\in A\\setminus B} d^A_i \\log(1 - \\pi^B_i)$, where $\\pi^B_i$ is the propensity of inclusion in the big dataset $B$, $A$ is the probability reference sample, and $d^A_i$ are its design weights. The first sum uses all big-data units directly; the second uses only reference units that are not in $B$, which is the key difference from previous methods that use all of $A$ for the second term. Its score equation sets a weighted sum of covariate vectors to zero and is solved by Newton-Raphson; the solution plugs into a logistic model $\\pi^B_i = \\exp(x_i^\\top \\theta)/(1 + \\exp(x_i^\\top \\theta))$ and then into the inverse-probability-weighted mean estimator. The mechanism that carries the argument is the removal of the overlap $A \\cap B$ from the non-inclusion term, which avoids double-counting the already observed big-data units and makes the estimating equation closer to the population likelihood.","core_discovery":"The central claim is that the Overlap-Excluded (OE) pseudo-likelihood estimator is a valid and efficient way to estimate selection propensities for a big non-probability dataset when a probability reference sample can be linked to it. OE maximizes a pseudo-log-likelihood that keeps the term over B for included units and uses only the reference units outside B (the set difference A \\ B), weighted by design weights, for the non-included term. Under a logistic propensity model, solving the OE score equation gives propensities that, when plugged into the inverse-probability-weighted mean estimator, yield unbiased population-mean estimates whenever the model is correct. In the paper's simulations with a population of 200,000, OE showed the lowest relative root mean squared error among the compared estimators at big dataset sizes of 50,000 and 140,000, and its 95% confidence intervals had coverage between 0.946 and 0.953. The authors further state that theoretical results, not included in the paper, confirm OE is at least as efficient as the CLW and KW estimators under Poisson sampling of the reference sample, and more efficient than the ALP estimator when propensities are large.","pith_inferences":["A design implication not pursued in the paper is that, if overlap membership is known, the reference sample could be deliberately drawn to maximize overlap with the big dataset, so that the A without B portion carries more information and the survey cost could be reduced.","The method's reliance on exact linkage suggests a natural stress test: corrupt a known fraction of overlap indicators and measure the bias; the paper does not quantify how sensitive OE is to linkage errors.","The same pseudo-likelihood construction could be extended to doubly robust estimation of totals and quantiles, using OE propensities as the fitted weights in a regression-assisted estimator.","If the reference sample uses a complex design with unequal inclusion probabilities, the plug-in variance may need an additional term; exploring this would make the method applicable beyond the simulation settings considered."],"forward_implications":["At big-data sizes of 50,000 and 140,000 in the simulation, the OE estimator had the lowest percentage relative root mean squared error of all compared estimators, with clear gains over the CLW, KW, WVL, and VD approaches.","When the propensity model is correctly specified, the OE-based IPW estimator is unbiased; this is claimed as a property of the method, not just an observed simulation outcome.","Coverage probabilities of 95% confidence intervals from the plug-in variance estimator stay close to nominal (0.946 to 0.953) across all simulated scenarios.","A theoretical comparison, reported but not derived in the paper, says OE is at least as efficient as CLW and KW under Poisson sampling of the reference sample, and more efficient than the ALP estimator when the true propensities are large."],"supporting_citations":[{"why":"Provides the CLW pseudo-likelihood estimator that OE is compared against and supplies the simulation setup used in Section 5.","marker":"Chen et al. (2020)"},{"why":"Establishes the data-integration approach with known overlap membership, which OE extends by excluding the overlap from the non-inclusion term.","marker":"Kim & Wang (2019)"},{"why":"Defines the ALP/WVL estimator, a direct comparator, and the relationship between indirect inclusion probability and propensity that OE is claimed to outperform at large propensities.","marker":"Wang et al. (2021)"},{"why":"Provides the VD pseudo-likelihood estimator, another baseline in the simulation study.","marker":"Valliant & Dever (2011)"},{"why":"Supplies the inverse-probability-weighted estimator form used in Section 4.1 and the review context for quasi-randomisation approaches.","marker":"Wu (2022)"}],"fun_headline_variants":["Overlap-excluded propensities cut big-data bias","Excluding overlap sharpens big-data estimate accuracy","Data integration yields efficient big-data propensities","OE estimator gives unbiased big-data population estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every unit in the reference sample A can be correctly classified as either present or absent in the big dataset B; if this linkage is wrong, the A without B set is misidentified and the OE score equation is misspecified, making the resulting propensity scores and IPW estimates biased.","fun_headline_variants_meta":{"raw":{"variants":["Overlap-excluded propensities cut big-data bias","Excluding overlap sharpens big-data estimate accuracy","Data integration yields efficient big-data propensities","OE estimator gives unbiased big-data population estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000494,"raw_usage":{"total_tokens":2408,"prompt_tokens":908,"completion_tokens":1500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":1440}},"tokens_in":524,"tokens_out":1500,"duration_ms":11253,"temperature":1.0,"reasoning_tokens":1440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:39:41.489492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the paper's simulation design, deliberately misclassify 5% of the overlap indicators in the reference sample and compare the OE IPW estimates with those from CLW under the same corrupted linkage; if OE's bias exceeds CLW's, the accurate-linkage assumption is the limiting premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CLW pseudo-likelihood estimator that OE is compared against and supplies the simulation setup used in Section 5."}],"review_version":1}