{"id":"3a5948ed-7f54-4d95-ac1b-aa0f0fd87692","arxiv_id":"2608.04919","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of data analysis and study design challenges in Long COVID research, covering definition, time-varying symptoms, and auxiliary-variable dependent sampling.","lead":"This paper reviews the statistical challenges of studying Long COVID, including how to define the condition without a gold standard and how to analyze data from studies that selectively test participants. It is a useful guide for biostatisticians working with the RECOVER cohort data and related observational studies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pseudo-label assumption is load-bearing and undefended; the LCRI's claim to rigorous definition depends on pre-infection symptoms not predicting infection, which is not established.","rationale":"The reader's weakest-assumption analysis correctly identifies the pseudo-label validity problem: training on infection history is only a valid surrogate for Long COVID if symptoms that distinguish infected from uninfected individuals are actually caused by the latent LC state. The paper's Figure 1 makes exactly this assumption (X not precursors of A; A affects X only through Y), but no evidence is given that the assumption holds, and the LCRI is then promoted as the only rigorous definition. This is a genuine soft spot. However, it is a soft spot in an illustrative definitional approach, not in the paper's central thesis that LC data have unusual statistical challenges. Those challenges—no gold standard, time-varying presentations, auxiliary-variable dependent sampling—are supported by the design descriptions and citations. The paper is a review, so the pseudo-label issue does not overturn its main message. The reader's conditional verdict is also driven by missing table/figure content and unsubstantiated evaluative language; my concern reinforces the need for a caveat about the LCRI's assumptions but does not change the overall verdict. I therefore recommend no change to the reader's conditional acceptance.","tokens_in":14985,"tokens_out":3706,"duration_ms":54793,"concrete_test":"Using RECOVER-Adult data, refit the LCRI Lasso with infection history as the outcome but add pre-infection baseline symptom burden or healthcare-utilization measures as covariates, then compare the resulting index with the original LCRI in an independent validation sample with clinically adjudicated LC outcomes. If pre-infection symptom burden materially changes the selected symptoms or coefficients, or if the LCRI's discrimination between adjudicated LC and non-LC infected controls attenuates once baseline symptoms are controlled, the pseudo-label assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 and Figure 1 present the negative-unlabeled strategy as a way to define Long COVID without a gold standard: infection history A is used as a pseudo-outcome for latent LC status Y, with symptoms X assumed not to be precursors of A and to be affected by A only through Y. This identifiability condition is load-bearing for the paper's strong claim that the resulting LCRI 'is the only strategy to our knowledge that rigorously defines LC' and minimizes misclassification of chronic conditions unrelated to SARS-CoV-2. The assumption is not defended, and it is questionable: pre-existing symptoms may predict infection risk through healthcare-seeking, testing access, or occupational exposure, so P(A|X) can differ substantially from P(Y|X). The text acknowledges false negatives from thresholding but does not address this false-positive mechanism or any validation of the pseudo-label premise beyond citing prior work. Because this construction underlies the paper's recommended 'specific' definition, the gap weakens the definitional discussion, even though the broader claim that LC research poses statistical challenges remains intact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a perspective/review article on statistical challenges in Long COVID (LC) research. It argues that LC studies face unique data analytic difficulties because there is no gold-standard definition, the condition has multiple and time-varying presentations, and sampling designs are often auxiliary-variable dependent. Section 2 discusses definitions, focusing on a negative-unlabeled approach that uses SARS-CoV-2 infection history as a pseudo-label to train a classifier, yielding the Long COVID Research Index (LCRI). Section 3 describes sub-phenotypes, waxing/waning symptoms, left censoring, and variant-related cohort composition. Section 4 describes two-phase and repeated auxiliary-variable dependent sampling as used in RECOVER-Adult and RECOVER-Pediatrics, and outlines selection bias, informative missingness, confounding, and loss to follow-up. The Discussion describes the NIH-funded Network of Biostatisticians for RECOVER (NBR). The central claim is that these features demand careful statistical design and analysis.","tokens_in":15115,"tokens_out":3672,"duration_ms":44920,"significance":"The paper performs a useful service by cataloguing and organizing the statistical challenges specific to LC research, and its descriptions of the RECOVER designs and LCRI appear consistent with the cited publications. Its main contribution is synthetic: it connects negative-unlabeled learning, two-phase sampling, and longitudinal missing-data concepts to a concrete, high-impact clinical domain, and it explicitly flags false-negative limitations of the LCRI and left censoring in contemporary cohorts. If the unsupported claims in Section 2 are appropriately qualified, the paper could serve as a practical orientation for statisticians entering this area. It is not a methodological development; its value lies in its accurate survey and its emphasis on design-aware analysis.","major_comments":[{"comment":"The negative-unlabeled strategy assumes that infection history A is a valid pseudo-outcome for latent LC status Y, and that symptoms X are not precursors of A and are affected by A only through Y. This identifiability condition is load-bearing for the claim that the LCRI 'rigorously defines LC' and minimizes misclassification of chronic conditions unrelated to SARS-CoV-2 infection, but it is not defended in this manuscript. Pre-existing symptoms may predict infection risk through healthcare-seeking behavior, testing access, or occupational exposure, in which case P(A|X) diverges from P(Y|X). The text acknowledges false negatives from thresholding but does not address this false-positive mechanism or provide independent validation of the pseudo-label premise beyond citing prior work. Please either provide evidence or a formal argument for this assumption, or substantially temper the claims about the LCRI.","section":"§2 and Figure 1"},{"comment":"The statement that the LCRI is 'the only strategy to our knowledge that rigorously defines LC' is not supported by a systematic comparison with other proposed definitions or by a formal criterion for what qualifies as 'rigorously defines.' Because the authors are also developers of the LCRI, this superlative claim needs either a concrete argument or a more modest formulation; as written, it reads as advocacy rather than assessment.","section":"Section 2"}],"minor_comments":[{"comment":"The caption contains the typo 'sstandard' and should be corrected to 'standard.'","section":"Figure 1 caption"},{"comment":"The text refers to the 'Omicon' variant; this should be 'Omicron.'","section":"Section 3"},{"comment":"The sentence beginning 'Resource-efficient cohort study designs an optimal subset of participants may be selected' is incomplete and should be reworded for clarity.","section":"Section 5"},{"comment":"The term 'na¨ıve' contains encoding artifacts; it should read 'naive.'","section":"Section 2 and reference list"},{"comment":"Table 2 is referenced in the Discussion but its content is not shown in the manuscript; please ensure the table is included in the final version.","section":"Discussion and Table 2"},{"comment":"The claim that 'individuals with no symptoms at all tend to be healthier than the general population' is an empirical assertion that would benefit from a supporting citation.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is heavily self-referential, promoting the LCRI and RECOVER designs developed by the authors, and the strong claim in Section 2 that LCRI is the only rigorous LC definition may be seen as advocacy rather than neutral review. The paper is a perspective piece rather than a methodological contribution; the editors should consider whether this fits the scope of a statistics journal and whether the unsupported pseudo-label assumption should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. This is a solid, clearly written review of the statistical headaches in Long COVID research, but it's not a methods paper, and its strongest claim—that the LCRI is the only rigorous definition of LC—rests on an assumption the paper never defends.\n\nThe useful parts: the paper lays out the definition problem honestly, walking through the sensitivity/specificity tradeoff and why neither a purely symptom-based definition nor the NASEM catch-all works well for research. The warning about using a raw symptom count as an outcome is practical and easy to overlook. The description of repeated auxiliary-variable dependent sampling in RECOVER-Adult is genuinely instructive, and the point about left censoring and structural intermittent missingness during reinfection is important and often missed. As a catalog of pitfalls for applied statisticians entering the field, this is valuable.\n\nThe soft spots. The pseudo-label construction—train on infection history as a proxy for latent LC—is the backbone of the LCRI. The paper states in Figure 1 that symptoms are not precursors of infection and that infection affects symptoms only through LC, but it never defends either condition. That's not a minor detail: pre-existing symptoms can influence testing behavior, healthcare access, and occupational exposure, so P(A|X) can be quite different from P(Y|X). The paper acknowledges false negatives from thresholding but is silent on the false-positive mechanism this assumption introduces. Given that, the claim that LCRI is 'the only strategy to our knowledge that rigorously defines LC' overreaches. It's the only strategy the authors have built, but rigorousness depends on an identifiability condition that is plausible but not established here. The paper is also heavily self-referential; pretty much all the main examples come from the authors' own RECOVER work. That's fine in a review, but the tone in the Discussion—NIH money, NBR, 'immense and novel scientific discoveries'—reads like a program announcement, not a methods review. And the arXiv version is missing the actual tables and figures, which makes it hard to fully evaluate.\n\nWho is it for? Applied statisticians and epidemiologists working on Long COVID or similar infection-associated chronic conditions. It's a good entry point, but someone wanting new methodology should go to the cited LCRI and two-phase sampling papers.\n\nI'd send it to peer review, but conditional on tempering the 'only strategy' claim, adding a paragraph that acknowledges the pseudo-label assumption and its failure modes, and providing the missing table/figure content. The core catalog is sound and worth publishing.","headline":"A clear, useful review of Long COVID statistical pitfalls, but its strongest claim about the LCRI rests on an undefended identifiability assumption, and the paper reads more like a program overview than a neutral methods survey.","tokens_in":15676,"tokens_out":3233,"would_cite":true,"duration_ms":35178,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62D05","62J07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Long COVID research requires statistical methods tailored to a condition with no gold-standard definition.","keywords":["Long COVID","negative-unlabeled data","auxiliary-variable dependent sampling","pseudo-label","LCRI","RECOVER","two-phase sampling","measurement error"],"falsifier":"A validation study that applies the LCRI to a population where infection history is known but Long COVID status is adjudicated by a blinded clinical panel, and shows that LCRI-positive never-infected individuals are common or that LCRI status fails to track clinically confirmed Long COVID trajectories, would falsify the pseudo-label core of the approach.","tokens_in":14758,"feed_emoji":"📊","tokens_out":4458,"duration_ms":48553,"temperature":0.7,"pith_summary":"This paper argues that Long COVID cannot be handled with off-the-shelf statistics because the condition lacks a gold-standard definition, changes over time, and presents differently across people. It lays out the resulting problems for study design and analysis: outcome misclassification, left-censored onset, and selection bias from sampling that depends on symptoms. The paper's central constructive claim is that a negative-unlabeled formulation, in which infection history serves as a pseudo-label, offers a workable route to defining Long COVID for research, exemplified by the Long COVID Research Index. A sympathetic reader would take away that valid Long COVID inference requires design-aware methods, not just larger cohorts.","feed_headline":"No gold standard means Long COVID statistics needs new rules","feed_subtitle":"Researchers must treat infection history as a stand-in label and symptom-triggered sampling as a source of bias.","key_machinery":"The central object is the negative-unlabeled data formulation of Long COVID status: never-infected individuals are 'negative' for Long COVID, while infected individuals are 'unlabeled' because they may or may not have the condition. This formulation licenses a pseudo-label classifier — Lasso-penalized logistic regression with infection history as the outcome and symptoms as predictors — whose nonzero coefficients form the Long COVID Research Index (LCRI), a weighted symptom score thresholded to define Long COVID. The other load-bearing mechanism is repeated auxiliary-variable dependent sampling, a two-phase design in which expensive tiered tests are administered to a subset selected on cheap auxiliary variables such as symptom triggers; the paper argues that ignoring this sampling mechanism induces selection bias that can attenuate or invert associations.","core_discovery":"The paper's central claim is that the defining features of Long COVID — absence of a gold standard, multiple sub-phenotypes, waxing and waning symptoms, and reliance on auxiliary-variable dependent sampling — create distinct statistical challenges that standard cohort analysis methods do not address. It argues that two moves make research possible: treating Long COVID status as negative-unlabeled data and using SARS-CoV-2 infection history as a pseudo-outcome to train classifiers such as the LCRI, and accounting for repeated auxiliary-variable dependent sampling in the analysis to avoid selection bias. The authors take the LCRI approach to be the only current strategy that rigorously defines Long COVID while minimizing misclassification of chronic conditions with other causes, and they describe structural intermittent missingness and left censoring as built into RECOVER-style designs.","pith_inferences":["If the pseudo-label assumption is wrong, the LCRI may measure general post-viral symptom burden rather than Long COVID specifically; this can be tested by external validation against clinician-adjudicated or biomarker-based diagnoses.","The same negative-unlabeled and auxiliary-variable dependent sampling framework could be transferred to other infection-associated chronic conditions, with a clear testable extension being cross-condition classifiers trained on symptom data from multiple post-infection syndromes.","The paper's review suggests that standardized reporting of sampling probabilities and trigger variables should become a required part of Long COVID study publications, a practice that would improve reproducibility across cohorts."],"forward_implications":["Studies that define Long COVID with a broad symptom-based or NASEM-style definition will tend to dilute effect sizes and lose power, so research definitions should be chosen for specificity.","Comparisons against individuals with an LCRI of 0 are preferable to comparisons against asymptomatics, because asymptomatics are healthier than the general population and bias results.","Analyses of tiered tests must account for the sampling mechanism and for informative missingness within the phase-two sample, or associations will be biased toward the null or otherwise distorted.","Left-censored enrollment and structural missingness around reinfections mean that time-to-recovery and trajectory analyses need methods that handle intermittent unobservability of the condition.","Symptom counts are poor outcomes because their distribution depends on the number of symptoms in the instrument and on correlations within organ systems, as demonstrated in the paper's comparison of naive counts with the LCRI."],"supporting_citations":[{"why":"Derives the adult Long COVID Research Index using Lasso-penalized logistic regression with infection history as the pseudo-outcome.","marker":"Thaweethai et al. (2023)"},{"why":"Develops penalized regression with negative-unlabeled data and shows the LCRI outperforms naive symptom counts in discriminative performance.","marker":"Reeder et al. (2025)"},{"why":"Provides the 2024 update of the RECOVER-Adult Long COVID Research Index.","marker":"Geng et al. (2024)"},{"why":"Supplies the RECOVER-Pediatrics study protocol and the derivation of pediatric LCRI using the same pseudo-label approach.","marker":"Gross et al. (2024)"},{"why":"Defines the NASEM characterization of Long COVID, the broad high-sensitivity definition the paper contrasts with the LCRI threshold.","marker":"Ely et al. (2024)"},{"why":"Surveys learning from positive and unlabeled data, providing the conceptual basis for the negative-unlabeled formulation.","marker":"Bekker and Davis (2020)"},{"why":"Establishes semiparametric regression methods for two-phase outcome-dependent sampling, the statistical foundation for auxiliary-variable dependent sampling designs.","marker":"Breslow et al. (2003)"},{"why":"Presents the RECOVER-Adult study protocol, the source of the repeated auxiliary-variable dependent sampling scheme with trigger variables.","marker":"Horwitz et al. (2023)"}],"fun_headline_variants":["Long COVID stats: no gold standard, new rules required","Long COVID research: pseudo-outcome replaces gold standard","Statistical workarounds for Long COVID's missing gold standard","Without a gold standard, Long COVID needs fresh statistics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that symptom patterns that separate people with a history of SARS-CoV-2 infection from people who were never infected also capture Long COVID itself; the paper adopts this pseudo-label assumption from earlier LCRI work without independently validating it.","fun_headline_variants_meta":{"raw":{"variants":["Long COVID stats: no gold standard, new rules required","Long COVID research: pseudo-outcome replaces gold standard","Statistical workarounds for Long COVID's missing gold standard","Without a gold standard, Long COVID needs fresh statistics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1258,"prompt_tokens":815,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":378}},"tokens_in":431,"tokens_out":443,"duration_ms":6122,"temperature":1.0,"reasoning_tokens":378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:40:57.239119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A validation study that applies the LCRI to a population where infection history is known but Long COVID status is adjudicated by a blinded clinical panel, and shows that LCRI-positive never-infected individuals are common or that LCRI status fails to track clinically confirmed Long COVID trajectories, would falsify the pseudo-label core of the approach.","supporting_citations":[{"cited_title":"and Thaweethai, Tanayott and Foulkes, Andrea S","cited_arxiv_id":null,"evidence_quote":"Develops penalized regression with negative-unlabeled data and shows the LCRI outperforms naive symptom counts in discriminative performance."},{"cited_title":"Wesley and Brown, Lisa M","cited_arxiv_id":null,"evidence_quote":"Defines the NASEM characterization of Long COVID, the broad high-sensitivity definition the paper contrasts with the LCRI threshold."},{"cited_title":"Large sample theory for semiparametric regression models with two-phase, outcome-dependent sampling , volume =","cited_arxiv_id":null,"evidence_quote":"Establishes semiparametric regression methods for two-phase outcome-dependent sampling, the statistical foundation for auxiliary-variable dependent sampling designs."},{"cited_title":"and Thaweethai, Tanayott and Brosnahan, Shari B","cited_arxiv_id":null,"evidence_quote":"Presents the RECOVER-Adult study protocol, the source of the repeated auxiliary-variable dependent sampling scheme with trigger variables."}],"review_version":1}