{"id":"3dcffc79-ada6-4679-9b5d-084c4b1cdf39","arxiv_id":"2608.04918","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The authors describe and illustrate inverse probability weighted estimators for outcome means and mean differences under two-phase auxiliary-variable-dependent sampling in the RECOVER Adult and Pediatric cohorts.","lead":"This paper presents inverse probability weighting estimators to correct for selective, auxiliary-variable-based testing in the RECOVER Long COVID cohort studies. It shows that accounting for the sampling design shrinks the estimated difference in abnormal smell test rates between Long COVID and non-Long COVID participants, from 11.2 to 6.3 percentage points.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The estimator relies on a monotone missingness assumption that discards post-return visits; if attendance is intermittent, the weights cannot repair the resulting selection bias, making the central bias-correction claim conditional on an untested structural assumption.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the imposed monotone missingness structure. This is not a peripheral detail but a structural change to the analysis sample. If intermittent attendance is realistic, the estimator discards valid test completions from returning participants, and the weights in equations (3), (4), and (6) do not account for the probability of return or for the selection induced by truncating the sample. The paper itself flags this as an open problem in the Discussion, but the central claim that the method corrects sampling bias is conditional on this assumption being satisfied or on the loss being ignorable. I considered other potential issues, such as the definition of the propensity score for time-varying exposure or the pooling of correlated visit-specific estimates, but those are either acknowledged in the text or represent standard IPW practice. The monotone missingness assumption is the most distinctive and least secured element: it is imposed 'to facilitate analysis' and not validated against the actual RECOVER visit process. A simulation with intermittent attendance is the natural check, and the absence of any simulation in the manuscript makes this concern concrete rather than speculative. Since the reader already reached CONDITIONAL with this as the weakest assumption, my assessment does not change the verdict.","tokens_in":13767,"tokens_out":8286,"duration_ms":107024,"concrete_test":"Run a simulation study with K=4 visits where visit attendance is intermittent: generate attendance from a model depending on time-varying symptoms and a latent outcome-related factor, with some participants missing one or more visits and returning later. Generate first-test completion and outcome conditional on attendance, then apply the paper's estimator after truncating to monotone attenders (discarding returns). Compare the estimated mean difference to the true pooled estimand and to an estimator that uses all observed visits. If the monotone-truncated estimate deviates by more than a pre-specified margin (e.g., 0.5 percentage points) while the all-visits estimator recovers the truth, the monotone assumption is not innocuous and the paper's central claim must be qualified to the monotone-attendance subpopulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the proposed inverse probability weights correct bias from auxiliary-variable-dependent sampling in RECOVER-style cohorts. That claim holds only under the monotone missingness structure imposed in Section 2.2: U_k=0 implies U_{k+1}=0, and at most one outcome per person. Under this structure, a participant who misses visit 3 but returns at visit 4 has U_4 forced to 0, so their visit-4 test completion is discarded even if it is their first eligible test. If the probability of missing a visit depends on auxiliary variables, exposure status, or the outcome itself—for example, a participant with no smell change is less motivated to attend—then the retained monotone sample is not representative of the eligible population. The inverse weights g^R_k and g^U_k are defined only on this truncated sample and do not model the probability of intermittent non-attendance or return. The Discussion explicitly defers 'potential loss of data due to monotonization' to future work, but the paper provides no simulation, sensitivity analysis, or argument that this loss is ignorable. Without such evidence, the estimator's bias-correction guarantee applies only to a subpopulation with monotone attendance, not to the full RECOVER cohort with potentially intermittent follow-up.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes inverse probability weighting estimators for outcome means and mean differences by exposure status in observational cohorts subject to auxiliary-variable-dependent (\"tiered\") sampling. Two designs are considered: repeated sampling over visits, as in RECOVER-Adult, and one-time sampling before exposure measurement and outcome collection, as in RECOVER-Pediatric. The estimators combine weights for visit attendance, sampling and test completion, and exposure propensity, with separate treatment of auxiliary variables used for sampling versus adjustment covariates. The adult estimator is illustrated with the UPSIT smell test, where the fully weighted analysis yields a Long COVID versus non-Long COVID difference of 6.3 percentage points (95% CI 1.3, 11.3) in severe microsmia or anosmia, compared with an unweighted difference of 11.2 points.","tokens_in":13985,"tokens_out":4455,"duration_ms":50461,"significance":"If the proposed estimators are valid, the paper addresses a practically important gap: tiered testing is common in large observational cohorts and EHR-based studies, yet accessible analytic guidance for such designs is limited. The paper's strengths include a clear formalization of the two sampling designs, an explicit distinction between auxiliary variables used for sampling and the adjustment covariate set, use of design-known sampling probabilities in the pediatric setting, and a concrete applied example with interpretable results. I found no circularity in the construction: the weights are estimated from auxiliary variables, exposure, and visit processes, not from the outcome. However, the central bias-correction claim is currently supported only by heuristic argument and one illustrative application; the paper provides no simulation, no formal consistency proof, and no sensitivity analysis for the structural monotone missingness assumption. These gaps are load-bearing for the method's general use.","major_comments":[{"comment":"The monotone missingness assumption (U_k=0 implies U_{k+1}=0, and at most one outcome per person) is load-bearing for the estimator, because the cumulative weights g^U_k and g^R_k are defined only for participants who remain continuously eligible without prior completion. Under intermittent attendance, a participant who misses visit 3 but returns at visit 4 has U_4 forced to 0, so their first eligible test at visit 4 is discarded even if it is their only eligible assessment; the inverse weights do not model the probability of return. The manuscript states that this structure was imposed \"to facilitate analysis\" and the Discussion defers \"potential loss of data due to monotonization\" to future work, but no simulation, sensitivity analysis, or empirical diagnostic is provided. The central bias-correction claim therefore currently holds only for a monotone-attendance subpopulation, not for the full RECOVER cohort if attendance is intermittent.","section":"Section 2.2, Eqs. (1)-(4)"},{"comment":"The estimator is proposed without a consistency theorem, regularity conditions, or a simulation study. The text states that sandwich variance formulae for M-estimators or bootstrap can be used, but it does not verify that these remain valid when the weights and the propensity score are estimated, nor does it validate the inverse-variance pooling in Eq. (7) when the variance estimates are themselves estimated. A formal M-estimation argument or a simulation study is needed to support the inferential claims, including the reported 95% confidence interval in Section 3.","section":"Section 2.2, Eq. (6) and the variance paragraph"},{"comment":"The text says that the model for π^A_k(a;W) must be fit \"inverse weighted by the probability defined by equation (3),\" but the exact fitting procedure is not specified: it is not stated whether this is a weighted logistic regression, how the weights are normalized, or how uncertainty in g^U_k propagates into the final estimator. Without this detail, Eq. (6) is not fully reproducible, and the claim that this step is \"necessary\" is not demonstrated. Please provide an explicit estimating-equation or algorithmic definition.","section":"Section 2.2, paragraph on estimating the propensity score"}],"minor_comments":[{"comment":"The word \"frequences\" appears twice in the example text and should be \"frequencies.\"","section":"Section 3"},{"comment":"The sentence \"Using estimates of each probability defined in equations (6-9)\" should refer to equations (8)-(11), since Eq. (6) is the adult estimator and the pediatric probabilities are defined in Eqs. (8)-(11).","section":"Section 2.3"},{"comment":"Table 2 is listed but its numerical contents are not included in the provided manuscript; the numbers quoted in the text are useful, but the table itself should be present in the final version.","section":"Table 2"},{"comment":"The in-text citation \"(Vaart, 1998)\" does not match the reference list entry \"A. W. van der Vaart\"; please correct the citation style.","section":"Section 2.2, variance paragraph"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of the journal and presents a useful practical framework. The central issue is that the bias-correction guarantee is conditional on a monotone missingness structure that is imposed rather than justified, and the paper provides no simulation or formal theory to support the estimator or its variance. These are fixable within the manuscript's scope by adding a simulation study, a sensitivity analysis for intermittent attendance, and a precise statement of the propensity-score fitting procedure. I see no citation-pattern concerns: the self-citations are to RECOVER study definitions and prior two-phase sampling methodology, not to the new estimator."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a competent application-and-guidance piece rather than a new method. The estimator in Eq. (6) and Eq. (12) is a direct combination of established IPW for two-phase sampling and propensity weighting, and the authors do not oversell it as more. What is genuinely useful is the careful mapping of RECOVER's repeated tiered-testing design into notation, the step-by-step implementation summary, and the concrete UPSIT illustration showing the weighted difference dropping from 11.2 to 6.3 percentage points. The writing is clear, the estimands are well defined, and the citation pattern is appropriate; the self-citations are to RECOVER study definitions rather than to the estimator itself. There is no circularity: the weighting models are fit without using the outcome means that the weights then produce.\n\nThe soft spots are real but not disqualifying. The load-bearing assumption is the monotone missingness structure in Section 2.2 (U_k=0 implies U_{k+1}=0, at most one outcome per person). If RECOVER participants attend intermittently, then post-return visits are discarded, and the inverse weights only model attendance among those who never missed a visit. The Discussion explicitly defers loss of data from monotonization to future work, but without a simulation or sensitivity analysis there is no evidence that this loss is ignorable. That is the main gap. The paper also provides no formal consistency proof, no simulation study, and no code or data; variance estimation is mentioned as sandwich or bootstrap but not validated. The numerical tables and figures are placeholders in this version, so the headline estimate cannot be independently checked.\n\nFor a methods paper, these gaps matter, but the central idea is not wrong. The paper does what it claims: it lays out a practical IPW strategy for a complicated real-world design and flags its own limitations honestly. It will be most useful to analysts working with RECOVER tiered-test data and to methodologists who want a clear statement of the design and a starting point for more robust estimators (e.g., doubly robust or monotone-missingness relaxations). With revision, it could be a solid applied-statistics contribution.\n\nI would send this to peer review, but with a request that the authors add a simulation or sensitivity analysis that relaxes monotone missingness and makes the variance estimation reproducible. My own verdict is conditional: the method is plausible and likely useful, but the paper's central claim is only as strong as the monotone missingness assumption it has not yet tested.","headline":"A useful, honest application of standard IPW to RECOVER's tiered-test design, whose main estimate rests on an untested monotone missingness assumption.","tokens_in":14579,"tokens_out":1700,"would_cite":false,"duration_ms":21681,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that inverse probability weighting over attendance, sampling, completion, and exposure propensity corrects the bias of auxiliary-variable-dependent 'tiered' testing, changing the estimated Long COVID smell-loss difference…","keywords":["Long COVID","inverse probability weighting","two-phase sampling","auxiliary variable dependent sampling","observational cohort","selection bias","missing data","RECOVER"],"falsifier":"Rerun the UPSIT analysis including all observed visits for participants who missed an eligible visit and later returned, fitting the same attendance, sampling, and completion models without the monotone restriction; if the weighted Long COVID versus non-Long COVID difference moves outside the reported 95% confidence interval of 1.3 to 11.3 percentage points, the monotone-missingness assumption is load-bearing. A simpler check is to compare the fully weighted estimate to a complete-case analysis restricted to participants who attended every eligible visit before first completion.","tokens_in":13542,"feed_emoji":"📊","tokens_out":10274,"duration_ms":102163,"temperature":0.7,"pith_summary":"The paper tries to establish that a single inverse-probability-weighting strategy can correct the selection bias introduced when observational cohorts administer expensive tier-2 tests to a subset of participants chosen according to auxiliary variables such as symptoms. It develops weighting recipes for two design variants: repeated sampling over visits with the exposure measured before sampling, as in RECOVER-Adult, and one-time sampling before the exposure is measured, as in RECOVER-Pediatrics. In the adult example, weighting for attendance, sampling, completion, and exposure propensity lowers the estimated prevalence of severe smell loss in both groups and reduces the Long COVID versus non-Long COVID difference from 11.2 to 6.3 percentage points (95% CI 1.3 to 11.3). The point is that ignoring the sampling mechanism can overstate how strongly Long COVID is tied to abnormal tiered-test results.","feed_headline":"Re-weighting shrinks Long COVID smell-loss gap to 6.3 points","feed_subtitle":"A new IPW adjustment for tiered sampling cuts the unweighted 11.2-point difference.","key_machinery":"The machinery is a product of inverse probabilities. For the adult design, each observed outcome at visit $k$ receives weight $1/(g_k^R \\pi_k^A)$, where $g_k^R$ is the modeled probability of first completing the tiered test at visit $k$ — combining visit attendance, eligibility, sampling, and completion under a monotone missingness structure ($U_k=0$ implies $U_{k+1}=0$, at most one outcome per person) — and $\\pi_k^A(a;W)=P(A_k=a|W)$ is the exposure propensity score defined for the entire cohort. Because the propensity score depends only on baseline covariates $W$ while attendance and sampling depend on auxiliary variables $Z$, the propensity model must be fitted with inverse weights $1/g_k^U$ so that the estimated propensity reflects the full cohort. Multiplying these inverses is what lets the estimator address loss to follow-up, the sampling design, differential completion, and confounding in one step.","core_discovery":"The paper's central claim is that for auxiliary-variable-dependent ('tiered') sampling, the adjusted exposure-specific mean of an outcome can be estimated consistently by weighting each observed outcome by the inverse of the joint probability of attending an eligible visit, being sampled and completing the test, and having the observed exposure level. In the repeated-sampling adult design, this weight is the inverse of the cumulative probability of first completing the test at visit k times the inverse of the exposure propensity score, where the propensity model is itself reweighted by the inverse probability of remaining on-study and eligible so that it reflects the full cohort rather than the tested subset. In the one-time pediatric design, the same logic multiplies the inverses of promotion, attendance, and completion probabilities by the inverse propensity weight. The visit-specific estimates are pooled by inverse-variance weighting, and variance is obtained from sandwich or bootstrap methods. Applied to the UPSIT smell test, the fully weighted estimate of the Long COVID versus non-Long COVID difference in severe microsmia or anosmia is 6.3 percentage points (95% CI 1.3 to 11.3), compared with 11.2 points unweighted.","pith_inferences":["If the monotone-missingness assumption fails, the reported 6.3-point difference is not guaranteed to be unbiased; a sensitivity analysis that restores intermittent attenders with visit-level indicators would show how much of the estimate depends on the recoding.","The same weighting structure could be transferred to other two-phase designs, such as immune-correlate analyses of vaccine trials, by plugging in design-known sampling probabilities for the promotion and attendance terms.","The gap between the unweighted 11.2-point and weighted 6.3-point difference is a direct quantification of how strongly trigger-driven oversampling inflated the crude smell-loss association, and a similar gap could serve as a diagnostic for other tiered tests.","A testable extension is to simulate full-cohort data from the fitted attendance, sampling, and completion models and compare the weighted estimator against the known truth, which would measure finite-sample bias and variance coverage."],"forward_implications":["Unweighted analyses of tiered tests in RECOVER-style cohorts will overstate outcome prevalence and group differences whenever sampling favors participants with symptoms or triggers; the weighted estimates provide the bias-corrected version.","The same product-of-inverse-probabilities recipe applies to EHR-based studies where test ordering depends on time-varying symptoms or laboratory results, because the data structure is the same even without a formal sampling scheme.","Pooled mean differences from this procedure are time-averaged associations over the follow-up period, not effects at a single time point, so they should be interpreted as averages across visits.","The approach covers concurrently measured exposure and outcome; exposure-to-outcome lagged questions, associations between two tiered tests, and trajectories of repeated tiered testing are stated by the authors as needing further work.","Because the trigger for one tiered test can depend on another tiered test (e.g., brain MRI triggered by UPSIT), the paper identifies those settings as requiring additional methods to avoid further selection bias."],"supporting_citations":[{"why":"Foundation of inverse-probability weighting for regression with missing data, which the paper extends to two-phase sampling.","marker":"(Robins et al., 1994)"},{"why":"Semiparametric theory for dependent censoring that motivates weighting for attendance and completion.","marker":"(Rotnitzky and Robins, 1995)"},{"why":"Framework for causal inference under outcome-dependent two-phase sampling that the proposed estimator adapts.","marker":"(Wang et al., 2009)"},{"why":"Large-sample theory for semiparametric regression with two-phase outcome-dependent sampling, underpinning the weighted estimation.","marker":"(Breslow et al., 2003)"},{"why":"Introduces the case-cohort design, the prototype of auxiliary-variable-dependent sampling the paper generalizes.","marker":"(Prentice, 1986)"},{"why":"RECOVER-Adult protocol defining the tiered sampling probabilities, eligibility criteria, and triggers used in the example.","marker":"(Horwitz et al., 2023)"},{"why":"Defines the 2024 Long COVID Research Index used to classify exposure status in the UPSIT example.","marker":"(Geng et al., 2025)"},{"why":"Provides comparison rates for smell/taste change and olfactory dysfunction in RECOVER-Adult used to interpret the weighted estimates.","marker":"(Horwitz et al., 2025)"}],"fun_headline_variants":["IPW cuts Long COVID smell-loss gap to 6.3 points","Weighted analysis trims Long COVID smell-loss gap to 6.3","Long COVID smell-loss difference drops to 6.3 points after IPW","Adjustment narrows Long COVID smell-loss gap from 11.2 to 6.3"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes missingness is monotone — once a person misses an eligible visit or completes the test, they never contribute again — so any skipped-and-returned attendance is recoded as dropout; if real attendance is intermittent, that recoding can itself create selection bias the weights do not repair.","fun_headline_variants_meta":{"raw":{"variants":["IPW cuts Long COVID smell-loss gap to 6.3 points","Weighted analysis trims Long COVID smell-loss gap to 6.3","Long COVID smell-loss difference drops to 6.3 points after IPW","Adjustment narrows Long COVID smell-loss gap from 11.2 to 6.3"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000666,"raw_usage":{"total_tokens":2998,"prompt_tokens":861,"completion_tokens":2137,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":2051}},"tokens_in":477,"tokens_out":2137,"duration_ms":19319,"temperature":1.0,"reasoning_tokens":2051,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:42:21.320203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the UPSIT analysis including all observed visits for participants who missed an eligible visit and later returned, fitting the same attendance, sampling, and completion models without the monotone restriction; if the weighted Long COVID versus non-Long COVID difference moves outside the reported 95% confidence interval of 1.3 to 11.3 percentage points, the monotone-missingness assumption is load-bearing. A simpler check is to compare the fully weighted estimate to a complete-case analysis restricted to participants who attended every eligible visit before first completion.","supporting_citations":[{"cited_title":"and Rotnitzky, Andrea and Zhao, Lue Ping , month = sep, year =","cited_arxiv_id":null,"evidence_quote":"Foundation of inverse-probability weighting for regression with missing data, which the paper extends to two-phase sampling."}],"review_version":1}