{"id":"c01d43ed-518f-4d9d-aeea-6d9921df8678","arxiv_id":"2508.00089","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Gradient boosting inside a two-step propensity-score weighting framework improves population estimates from nonprobability samples when selection is nonlinear, but not uniformly over all settings.","lead":"This paper replaces logistic regression with gradient-boosted trees in the first step of a two-step weighting method for nonprobability samples. The boosted version often reduced bias and variance in simulations and in a health survey application, but simpler logistic methods still did better in simple settings.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 contradicts the 'consistently outperforms' claim: Boost2PS has larger MSE than 2PS in all four mild scenarios (e.g., 3.69 vs 2.03 in I0Q0), and no consistency result justifies extrapolating the favorable complex-scenario results.","rationale":"The reader identifies the missing theory for the GBM-plus-logistic additive balancing score as the weakest assumption; I agree that this is a real gap. My stress-test adds that the paper's own Table 1 makes the gap concrete: in the four mild scenarios the GBM-based score is substantially worse than logistic 2PS, so it is not merely that a theorem is absent, but that the proposed combination demonstrably fails to balance when the true mechanism is close to linear. That internal evidence is more decisive than the absence of theory alone, because it shows the favorable results are confined to scenarios where the parametric comparator is misspecified by construction. The real-data example treats weighted NHIS as a benchmark and reports point comparisons without formal tests of differences between methods, so it does not independently establish general superiority. The paper still has coherent simulation evidence and a plausible practical contribution, so the appropriate verdict remains conditional: either provide a consistency argument or a finite-population balance verification, narrow the 'consistently outperforms' claim to complex misspecification settings, and release code or detailed replication instructions. The lack of released code and the 'Data Availability: N/A' statement further weaken the ability to confirm the results, but they are secondary to the balancing-score validity concern. No ad hominem is intended; the issue is with the strength of the evidence, not the authors' conduct.","tokens_in":21332,"tokens_out":14279,"duration_ms":156680,"concrete_test":"Re-run the Section 4 simulations and, for each replication and scenario, compute the finite-population balance of the final Boost2PS weights: the maximum absolute standardized difference between the weighted covariate means in the pseudo-weighted nonprobability sample and the true finite-population means, and the same quantity for oracle weights w_i = 1/pi_i^(c). If in Scenarios 5-8 Boost2PS's finite-population balance is not within Monte Carlo error of the oracle balance, the additive score is not acting as a valid balancing score and the observed bias reduction is not explained. Also report paired differences in MSE between Boost2PS and 2PS with Monte Carlo standard errors to test whether the 'consistently outperforms' wording survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The method's validity rests on the claim that the additive score b(x) = b1_GBM(x) + b2(x; gamma-hat) is a balancing score for the nonprobability sample relative to the finite population, so that pseudo-weighted covariate means match the finite population. This is the assertion that would have to be true for the MSE reductions to generalize. The paper does not establish it. Li (2024) proves the 2PS construction for two parametric logistic scores sharing g(x); here b1 is a tuned GBM log-odds, and no consistency, convergence-rate, or double-robustness result is given for the combination. The only balance evidence is ASMD comparisons against the reference survey (Table C), not against the true finite population; in the simulations, the paper reports outcome bias and MSE but not a direct finite-population balance check. The internal Table 1 also shows the balancing property fails in simple settings: Boost2PS MSE (x10^4) is 3.69, 2.56, 3.09, 3.00 vs 2PS 2.03, 2.12, 2.19, 2.15 in Scenarios 1-4. Consequently, the favorable Scenarios 5-8 and NHANES III results are not backed by a mechanism verified outside the specific simulation design. This is addressable by adding a direct balance check and either a theoretical statement or a narrowed claim, but as written the strongest claim is too broad.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes two pseudo-weighting estimators, Boost2PS and Boost1PS, that replace logistic propensity models with gradient boosting in Li's two-step pseudo-weighting framework. In Boost2PS, a GBM log-odds balancing score from the first step is added to a logistic balancing score from the second step, and pseudo-weights are constructed as exp{−b1−b2}; hyperparameters are tuned by minimizing the averaged absolute standardized mean difference (ASMD). The methods are evaluated in eight simulation scenarios with varying degrees of nonlinearity and non-additivity, and in a real-data application using NHANES III as a nonprobability sample and the 1994 NHIS as the reference survey, with mortality outcomes as benchmarks. The paper's central claim is that Boost2PS consistently outperforms the original 2PS method, especially under moderate to severe nonlinearity.","tokens_in":21668,"tokens_out":5693,"duration_ms":55024,"significance":"The application is timely and the empirical strategy is coherent: the simulations span linear to severely nonlinear selection mechanisms, the real-data benchmark is external (weighted NHIS estimates), and outcome variables are not used in weight construction, so the evaluation is not outcome-circular. The proposed use of gradient boosting is practical, builds on existing R packages (gbm and twang), and the real-data example provides a concrete illustration. However, the theoretical justification is incomplete and the headline claim overstates the evidence, because the paper's own Table 1 shows that Boost2PS has larger MSE than 2PS in the four mild scenarios. If the authors add a balancing-score consistency result or clearly narrow the claim to misspecified/complex settings, the method would be a useful contribution to the nonprobability-sample literature.","major_comments":[{"comment":"The abstract and Section 1 state that Boost2PS 'consistently outperforms' the original 2PS method, but Table 1 shows the opposite in Scenarios 1-4: the MSE (×10^4) for Boost2PS versus 2PS is 3.69 versus 2.03 in I0Q0, 2.56 versus 2.12 in I0Q1, 3.09 versus 2.19 in I1Q0, and 3.00 versus 2.15 in I1Q1. Section 4.4 itself acknowledges that in simpler scenarios the parametric methods performed better and were approximately unbiased. The consistent-superiority claim is therefore contradicted by the manuscript's own evidence; the claim should be restricted to moderate/severe nonlinearity or non-additivity, or a formal decision criterion over the scenario space should be used.","section":"Section 1 and Section 4.4 / Table 1"},{"comment":"The final adaptive balancing score b(x;θ,γ̂)=b̂1^T(x;θ)+b2(x;γ̂) and the pseudo-weights exp{−b̂1−b2} are introduced without a theorem. Li (2024) justifies the 2PS construction for two parametric logistic balancing scores that share the same covariate function g(x); no consistency, convergence-rate, or balancing-score property is established for a tuned GBM log-odds combined with a logistic balancing score. Since the authors attribute the simulation gains and the real-data bias reductions to this construction, the manuscript should either prove under appropriate regularity conditions that this additive score is a balancing score for the nonprobability sample relative to the finite population, or explicitly state that the method is heuristic and only empirically validated. Without this, the proposed mechanism is not established outside the specific simulation design.","section":"Section 3.2"},{"comment":"Hyperparameters are selected by minimizing ASMD (equation 3.3), and the same ASMD measure is subsequently used as the covariate-balance diagnostic in Table C. This makes the balance evidence partly circular: a method tuned to minimize ASMD on the same sample will tend to show small ASMD by construction. The simulations report outcome bias and MSE but do not report direct finite-population covariate balance after weighting, which is the mechanism the method claims to control. Please report a direct balance check (e.g., absolute standardized differences between the pseudo-weighted nonprobability sample and the true finite population in the simulations, or an out-of-sample ASMD) and use an independent criterion or cross-validation for hyperparameter selection, or at least discuss the potential overfitting of the balance criterion.","section":"Sections 3.3, 4.3, and Appendix C"}],"minor_comments":[{"comment":"The text says to use 𝑤̂i^{Boost1PS} to replace 𝑤̂i^{Boost2PS} in formula (3.4), but equation (3.4) is the weighted loss function; the estimator formula is (3.2).","section":"Section 3.5"},{"comment":"The text says '10 base covariates (V1,...,V7)', but only seven base covariates are listed; please correct the count or define the missing covariates.","section":"Section 4.1"},{"comment":"Figure 5 and its caption compare only Naïve, Boost1PS, and Boost2PS, while the surrounding discussion refers to four pseudo-weighting methods and claims an advantage over 2PS; including 2PS in the standard-error comparison would make the claimed advantage verifiable.","section":"Figure 5 and Section 5"},{"comment":"For a methods paper, releasing simulation code and the data-processing code for the real-data example would substantially aid reproducibility; the current 'Data Availability: N/A' entry is a limitation.","section":"Data Availability"},{"comment":"There are several typographical and formatting issues, such as 'tunning' for 'tuning' in Sections 3.1 and 3.3 and inconsistent spacing in 'Boost 2PS' versus 'Boost2PS'; a careful copyedit is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's contribution is incremental but potentially useful, and the real-data example is informative. However, the theoretical gap concerning the additive GBM-plus-logistic balancing score and the overbroad consistency claim are load-bearing issues that should be addressed before publication. I would encourage a major revision rather than rejection, because the central empirical strategy is coherent and the paper's own results are informative; the authors can likely fix the issues by adding a direct balance check and narrowing the claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does something simple and mostly useful: it replaces the logistic first step in Li's 2PS pseudo-weighting with a gradient-boosted log-odds balancing score, and adds a weighted-loss variant for the one-step estimator. The specific Boost2PS estimator is new relative to what the paper cites; it's a direct substitution inside Li's framework, but a sensible one. The simulation study is well-designed, separating additive/nonlinear and interactive structures across eight scenarios, and the NHANES III–NHIS example is a fair, relevant test bed. The balance tables and variance-ratio diagnostics are a real strength.\n\nThe soft spots are not fatal but they are real. First, the paper overclaims. The abstract and introduction say Boost2PS 'consistently outperforms' 2PS, but Table 1 shows the opposite in the four mild scenarios: MSE for Boost2PS is 3.69 vs 2.03 for 2PS in Scenario 1, and similarly worse in Scenarios 2–4. The results text quietly concedes this, so the headline needs to be narrowed to 'moderate to severe nonlinearity.' Second, there is no theory for the additive combination of a GBM score and a logistic score. Li's 2024 result covers parametric logistic balancing scores; the tree-based version has no consistency or double-robustness guarantee, so the simulation gains are empirical. That is acceptable for an applied methods paper, but it should be said plainly and a direct balance check against the true finite population in the simulations would help. Third, hyperparameters are tuned by minimizing ASMD, which is also the criterion used to report balance, so there is a mild circularity; the bias/MSE results are less affected, but the balance numbers are optimistic. Finally, no code or data are released, which is a real omission for a methods paper.\n\nOverall, the central idea is plausible and the evidence in the complex scenarios supports it. The paper deserves peer review. Send it out, let a referee push on the overclaim and the missing theory, and require code or a detailed tuning sensitivity analysis before publication. I'd bring it to a reading group; it will generate a good discussion about when ML weights actually help.","headline":"A practical, well-tested GBM variant of Li's 2PS that helps in complex scenarios, but the paper overstates the case and lacks theory for the boosted score.","tokens_in":22183,"tokens_out":2916,"would_cite":true,"duration_ms":25899,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62G08"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that replacing the logistic regression in the first step of the two-step pseudo-weighting procedure with gradient boosting reduces bias and mean squared error of population mean estimates when nonprobability sample…","keywords":["gradient boosting","nonprobability samples","pseudo-weighting","propensity score","selection bias","population mean estimation","two-step weighting","NHANES III"],"falsifier":"Simulate a nonprobability sample whose selection mechanism is a smooth nonlinear function that shallow trees approximate poorly (e.g., participation probability proportional to exp(sin(x1)+$x2^{3}$)), run Boost2PS with the paper's ASMD-based tuning, and compare the pseudo-weighted mean to the known population mean; if the estimator remains substantially biased when the same data is fit well by a correctly specified logistic model, the paper's central claim about GBM's flexibility would be falsified.","tokens_in":21096,"feed_emoji":"📊","tokens_out":6597,"duration_ms":59612,"temperature":0.7,"pith_summary":"This paper seeks to establish that gradient boosting can replace logistic regression in the first step of Li's two-step propensity-score pseudo-weighting (2PS) method and yield more accurate population mean estimates from nonprobability samples. The authors propose Boost2PS, which estimates the first-step balancing score with a gradient boosting machine and adds it to the second-step logistic balancing score, producing pseudo-weights of the form exp(-b1 - b2). Monte Carlo simulations with eight selection-mechanism scenarios show that Boost2PS consistently achieves the lowest absolute relative bias under moderate to severe nonlinearity or non-additivity, particularly in the most complex scenario, while also reducing empirical variance and mean squared error compared with parametric 1PS and 2PS. In an application using NHANES III as a nonprobability sample and NHIS as the reference survey, Boost2PS produced the smallest bias against NHIS benchmarks for several mortality outcomes and reduced bootstrap standard errors. If these results hold, the method gives survey researchers a practical way to hedge against propensity-model misspecification when estimating population quantities from volunteer samples.","feed_headline":"Boosted pseudo-weights cut bias in nonprobability-sample estimates","feed_subtitle":"Swapping logistic regression for gradient boosting in two-step weighting lowers bias when selection is nonlinear.","key_machinery":"The central object is the GBM-estimated balancing score b1(x), obtained by iteratively adding regression trees to minimize the log-loss for membership in the nonprobability sample versus the unweighted probability sample. Because GBM works on the logit scale, the resulting b1(x) is on the same scale as the logistic balancing score b2(x) from the second step, so the two can be added to form the final balancing score b(x) = b1(x) + b2(x). The pseudo-weight is then exp(-b1 - b2), and the hyperparameters (shrinkage, tree number, interaction depth) are selected to minimize the average absolute standardized mean difference between the weighted nonprobability sample and the unweighted probability sample. This additive-score construction is what lets a flexible tree ensemble plug directly into Li's two-step framework.","core_discovery":"The central claim is that the two-step pseudo-weighting estimator remains valid, and becomes more accurate, when the first-step balancing score is learned by gradient boosting rather than by a parametric logistic model. Because GBM models the log-odds of membership directly, its output can be added to the logistic balancing score from the second step to form the final score b(x) = b1(x) + b2(x), with pseudo-weights exp(-b1 - b2). The paper reports that in simulations with nonlinear and non-additive participation mechanisms (Scenarios 5, 7, and 8), Boost2PS consistently achieves the lowest absolute relative bias, with the largest gains in the severely nonlinear Scenario 8, and that its empirical variance and mean squared error are stable and often the smallest among the four methods compared. In the real-data analysis, boosting improves covariate balance between the pseudo-weighted NHANES III sample and the sample-weighted NHIS, and Boost2PS estimates of mortality prevalence are closest to the NHIS benchmarks for most outcomes.","pith_inferences":["A natural extension the paper does not test is whether the second-step logistic adjustment could itself be replaced by a nonparametric score; if the additive structure is what matters, the gains might persist with two flexible steps.","The ASMD tuning criterion targets covariate balance only; weighting covariates by their outcome-association strength (which the paper mentions as an option) could further reduce bias for a specific outcome like diabetes mortality, where all methods underperformed.","The absence of a consistency theorem for the tree-based balancing score means the method's reliability outside the simulated scenarios is an open empirical question; a cross-validation study on additional real nonprobability and reference survey pairs would test the claim's generality.","If the approach transfers to other balancing-score estimators, it could offer a general recipe for injecting nonparametric flexibility into design-based weighting without abandoning the reference-survey framework."],"forward_implications":["Under the paper's simulations, researchers analyzing nonprobability samples with complex self-selection mechanisms can expect smaller bias from Boost2PS than from logistic 1PS or 2PS, without needing to know the correct functional form in advance.","The bootstrap variance estimator proposed for Boost2PS is slightly conservative (variance ratios roughly 1.15 to 1.27), so uncertainty statements built on it will not understate sampling variability.","In the NHANES III / NHIS illustration, Boost2PS yields the pseudo-weighted covariate distributions closest to the reference survey among the methods considered, which supports its use in health-outcome prevalence estimation.","Because the method is implemented with existing R packages (gbm, twang, survey), the proposed gains are available without new software."],"supporting_citations":[{"why":"Supplies the two-step pseudo-weighting framework and the theory for the parametric balancing-score combination that Boost2PS extends.","marker":"Li (2024)"},{"why":"Defines the one-step adjusted logistic pseudo-weighting estimator that serves as the parametric comparator and whose real-data example is reused.","marker":"Wang et al. (2021)"},{"why":"Establishes gradient-boosted regression as a propensity score estimator that produces log-odds balancing scores, the basis of the new first step.","marker":"McCaffrey et al. (2004)"},{"why":"Provides the gradient boosting algorithm that defines the iterative tree ensemble and shrinkage used for the balancing score.","marker":"Friedman (2001)"},{"why":"Provides the twang/gbm software used to fit the boosted models and compute the ASMD balance criterion.","marker":"Ridgeway et al. (2013)"},{"why":"Documents boosted kernel weighting for nonprobability samples, providing prior use of boosting in this setting that motivates the approach.","marker":"Kern et al. (2021)"}],"fun_headline_variants":["Boosted pseudo-weights cut bias in nonprobability samples","Nonlinear selection? Boosted weights keep estimates on target","Two-step pseudo-weighting gets a gradient boosting boost","GBM pseudo-weights reduce bias in skewed survey data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method rests on the assumption that a gradient-boosted log-odds score added to a logistic balancing score yields pseudo-weights that genuinely balance the volunteer sample against the target population; if that additive combination fails to balance, the bias reductions seen in simulations are not guaranteed elsewhere.","fun_headline_variants_meta":{"raw":{"variants":["Boosted pseudo-weights cut bias in nonprobability samples","Nonlinear selection? Boosted weights keep estimates on target","Two-step pseudo-weighting gets a gradient boosting boost","GBM pseudo-weights reduce bias in skewed survey data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000734,"raw_usage":{"total_tokens":3309,"prompt_tokens":997,"completion_tokens":2312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":2246}},"tokens_in":613,"tokens_out":2312,"duration_ms":16565,"temperature":1.0,"reasoning_tokens":2246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:23:09.994576+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a nonprobability sample whose selection mechanism is a smooth nonlinear function that shallow trees approximate poorly (e.g., participation probability proportional to exp(sin(x1)+$x2^{3}$)), run Boost2PS with the paper's ASMD-based tuning, and compare the pseudo-weighted mean to the known population mean; if the estimator remains substantially biased when the same data is fit well by a correctly specified logistic model, the paper's central claim about GBM's flexibility would be falsified.","supporting_citations":[{"cited_title":"nonprobability sample","cited_arxiv_id":null,"evidence_quote":"Defines the one-step adjusted logistic pseudo-weighting estimator that serves as the parametric comparator and whose real-data example is reused."}],"review_version":1}