{"id":"aa16e881-4605-4137-8350-a7415c6d2c6d","arxiv_id":"2504.12415","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Pregnancy cohorts built from observed deliveries or outcomes are biased when outcomes are missing, while cohorts including all pregnancies are unbiased only when missingness is due to measured covariates.","lead":"This paper uses simulated data on 10 million pregnancies to compare two ways of identifying pregnancies in health care databases and measures how each method biases estimates of prenatal drug effects. It finds that including all pregnancies, not just those with recorded outcomes, removes the bias only when the missing outcomes are explained by measured patient characteristics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prenatal approach's claimed advantage depends on the idealized assumption of a probability-1 prenatal encounter at exactly 9 weeks; variable real-world encounter timing could alter the relative performance.","rationale":"The paper is a carefully specified simulation with a transparent data-generating process and a true-value benchmark from potential outcomes. Its internal logic is sound: outcome-based identification conditions on post-exposure events and can induce selection bias, while the prenatal approach with all pregnancies observed can recover the target effect when missingness is explainable by measured covariates. The reader's conditional verdict is appropriate because of the acknowledged simplification. I agree with the reader's weakest-assumption identification: the probability-1 encounter at exactly 9 weeks is the condition that makes the prenatal approach's advantage hold. This is the single most load-bearing external-validity concern because the title and discussion generalize to real-world data, where this assumption is known to fail. Without it, the observed-pregnancies sample is not necessarily the full target population, and the demonstrated unbiasedness under MAR may not persist. The check I propose directly tests whether the central conclusion survives relaxation of this assumption. If it does, the paper's claim is robust; if not, the conclusion should be recast as conditional on universal first-trimester capture. The paper already acknowledges the limitation, so I do not think the verdict should change from CONDITIONAL; it remains the right call pending this sensitivity analysis.","tokens_in":42401,"tokens_out":8266,"duration_ms":99190,"concrete_test":"Re-run the simulation with variable first prenatal encounter timing, e.g., drawing first-encounter week from a realistic distribution (weeks 6-20, with a small fraction having no prenatal encounter), while keeping all other data-generating mechanisms, missingness mechanisms, and analysis methods identical. If the observed-pregnancies estimates under 100% MAR remain unbiased after standardizing to the captured pregnancies, the concern is resolved; if bias appears, the prenatal approach's advantage is an artifact of universal 9-week capture and the paper's real-world conclusion would need to be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison rests on the simulation's assumption, stated in Methods S6, that every pregnancy has a prenatal encounter with probability 1 at exactly 9 weeks gestation. This guarantees that the 'prenatal approach' identifies the complete target population at a common time zero, so the observed-pregnancies analytic sample is the full target population before any missingness is induced. Under that condition, the finding that this sample is unbiased when all missingness is due to measured covariates (100% MAR) follows from standard ignorability of censoring within severity-rurality strata. In real claims and EHR data, first-encounter timing is variable, some pregnancies may never be captured by a prenatal visit, and care-seeking may be correlated with hypertension severity, rurality, and outcome risk. If the probability-1 common-time-zero encounter fails, the observed-pregnancies sample is itself selected: pregnancies with late or no prenatal care are missing entirely, not merely missing an outcome. The standardization to the analytic sample's covariate distribution would then not recover the target-population effect, and the prenatal approach could be biased even under MAR. The authors explicitly acknowledge this simplification in the Limitations section, noting that it 'does not reflect the substantial variability around that timing,' but the quantitative conclusion that only the observed-pregnancies sample estimates absolute risks without bias under MAR is conditional on this idealized assumption. Because the paper's title and framing concern real-world data, this is the most load-bearing threat to the generalizability of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies bias in estimates of prenatal treatment effects arising from how pregnancies are identified in real-world healthcare data. Using a simulation of 10,000,000 pregnancies under a hypothetical trial of antihypertensive initiation, the authors generate potential outcomes for preeclampsia and miscarriage, induce missingness under MAR (due to measured severity/rurality), MNAR (due to unobserved miscarriage), and mixtures, and construct three analytic samples: observed deliveries, observed outcomes, and all observed pregnancies (the prenatal approach). Treatment effects are estimated by nonparametric direct standardization, or by Aalen-Johansen estimation for the all-pregnancies sample, and bias is computed against true values from the simulated potential outcomes. The central findings are that under pure MNAR all three analytic samples are similarly biased, whereas under pure MAR only the all-pregnancies sample recovers unbiased absolute risks and effect estimates; the observed deliveries and observed outcomes samples remain biased. The paper frames outcome-based pregnancy identification as selection bias and the prenatal approach as a missing-data problem, and argues the prenatal approach enables sensitivity analyses and bounds.","tokens_in":42618,"tokens_out":6194,"duration_ms":70865,"significance":"If the results hold, the paper fills a genuine gap: prior work has focused on conditioning on live birth, but not on the bias induced by the pregnancy identification strategy itself. The simulation is carefully constructed, large, and checked against known true values from potential outcomes rather than fitted to a target, and the code is publicly available. The comparison of observed deliveries versus observed outcomes versus all pregnancies under MAR/MNAR gives a concrete, useful message for pharmacoepidemiology: outcome-based identification cannot generally be fixed by analytic methods, while the prenatal approach only removes bias when missingness is explained by measured covariates. The paper also contributes a clear DAG-based taxonomy and demonstrates full sensitivity bounds for the prenatal approach. These strengths make the study a solid methodological contribution, provided the central assumption discussed below is either relaxed in additional simulations or its conditioning role is stated more sharply.","major_comments":[{"comment":"The central quantitative comparison is conditional on the assumption, stated in Methods S6, that every pregnancy has a prenatal encounter with probability 1 at exactly 9 weeks gestation. This guarantees that the 'observed pregnancies' analytic sample is the full target population at a common time zero, so unbiasedness under 100% MAR follows from correct specification of the censoring model within severity-rurality strata. If first-encounter timing varies or some pregnancies are never captured by a prenatal visit, the observed-pregnancies sample is itself selected and standardization to its covariate distribution will not recover the target-population effect even under MAR. The manuscript acknowledges this limitation but does not quantify how the relative performance of the three approaches changes when the probability-1 common-time-zero assumption fails. Because the paper's practical message ('only among all pregnancies did bias decrease as the proportion of missingness due to measured variables increased') depends on this assumption, I ask the authors to add a simulation arm that varies first-encounter timing or capture probability, or to explicitly and prominently restrict the conclusion to the idealized setting and state that the ordering of approaches is not established when capture is incomplete.","section":"Methods S6 and Discussion, Limitations"}],"minor_comments":[{"comment":"The caption labels the fourth panel '(D) MAR', but panels are already labeled (A), (B), (C), (D) for MNAR, mixed, and MAR in the main text; this should read '(F) MAR' to match the corresponding panels in Figure 4 and the text.","section":"Figure 5 caption"},{"comment":"In the final block of Table S10 ('Initiation increases the risk of abortion and does not affect the risk of preeclampsia'), the numbers of pregnancies in the target population appear to be about 2.5 million per arm rather than the 5 million per arm used in all other scenarios; this is inconsistent with the stated 10 million pregnancies and should be corrected or explained.","section":"Table S10"},{"comment":"The supplemental tables use the term 'abortion' where the main text uses 'miscarriage'; the terminology should be harmonized to avoid confusion, since 'abortion' has a different clinical meaning.","section":"Tables S4-S10"},{"comment":"The rationale for assigning a follow-up time of 0.0001 to pregnancies with only the initial prenatal encounter should be stated more explicitly; this effectively censors those pregnancies at baseline, and the text should say so and note that the results are insensitive to the choice of this small constant.","section":"Methods S10"},{"comment":"The main-text statement 'All code is available on GitHub ([LINK TO FOLLOW UNBLINDING])' contains a placeholder; the actual repository URL appears in the front matter and should be inserted in the body of the manuscript.","section":"Data and Code Availability"},{"comment":"The third panel of Figure S8 labels the scenario 'All MNAR: 100% Severity, 0% Miscarriage'; since this panel represents all missingness due to measured severity/rurality, it should be labeled 'All MAR' to match the other figures.","section":"Figure S8 panel headers"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed simulation paper with reproducible code and a clear message. My main concern is that the idealized probability-1 fixed-time prenatal encounter is load-bearing for the paper's practical conclusion; a targeted sensitivity analysis or a sharper restriction of the claim should be achievable without changing the paper's scope. I do not see any indication of methodological circularity or integrity problems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful, transparent simulation study that makes a real and useful point—switching from outcome-based to prenatal pregnancy identification does not eliminate bias from missing pregnancy outcomes; it only gives you a basis for dealing with the bias when missingness is explained by measured covariates. The comparison across 36 scenarios is a genuine extension of prior live-birth conditioning work, and the code is public.\n\nWhat the paper does well: it defines the three analytic samples clearly, uses a multinomial competing-events outcome, and benchmarks every estimator against true potential-outcome values from 10 million simulated pregnancies. The result that observed deliveries and observed outcomes can be similarly biased, and that risk ratios can be worse than risk differences, is practically useful. The bounds analysis for the observed-pregnancies sample is also a good addition. The authors are upfront about many simplifications.\n\nThe soft spots: the cleanest headline result—that the observed pregnancies analytic sample estimates absolute risks without bias under 100% MAR—depends on the assumption that every pregnancy has a probability-1 prenatal encounter at exactly 9 weeks. The authors acknowledge this in the limitations, but it is load-bearing. In real claims/EHR data, first-encounter timing varies and some pregnancies never appear; then the prenatal sample is itself selected, and MAR may not be recoverable by standardizing on measured covariates. I do not think this kills the paper’s main qualitative lesson, but it should be tested or at least stated more prominently as a boundary condition. A sensitivity analysis with variable or missing first-encounter timing would materially strengthen it.\n\nMinor issues: Methods S10 assigns 0.0001 time to pregnancies with only the initial encounter; that is a modeling shortcut worth a sentence of justification. The absence of Monte Carlo uncertainty and reported seeds is not a major problem at this sample size but would be nice for reproducibility. The GitHub link is present; the manuscript should confirm a commit hash.\n\nOverall: the central claim holds up under the stated assumptions, and the paper is honest about where those assumptions are idealized. It deserves serious peer review. I would send it to a statistician with missing-data/competing-risks experience and ask for the encounter-timing sensitivity analysis; otherwise I expect it to be a useful methods paper for perinatal pharmacoepidemiology.","headline":"A transparent, well-benchmarked simulation showing that outcome-based pregnancy identification biases effect estimates and that the prenatal approach only fixes the problem when missingness is explained by measured covariates; the cleanest unbiasedness result is conditional on an idealized probability-1 prenatal encounter at 9 weeks.","tokens_in":43218,"tokens_out":3146,"would_cite":true,"duration_ms":36335,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The way researchers identify pregnancies in healthcare data can bias estimates of prenatal drug effects; the bias is only removable when missing outcomes are caused by measured variables.","keywords":["pregnancy identification","selection bias","missing not at random","prenatal exposure","miscarriage","pharmacoepidemiology","competing events","missing data"],"falsifier":"Analyze a real claims or electronic health record database linked to a complete pregnancy registry that records miscarriages and deliveries for the same source population, estimate a prenatal drug effect three ways (observed deliveries, observed outcomes, observed pregnancies with standardization), and check whether the observed-pregnancies estimate becomes unbiased when all measured predictors of missingness are adjusted for; if it remains biased, the central claim is falsified. A cheaper check is analytic: modify the simulation so first prenatal visits occur at varying gestational ages rather than exactly 9 weeks and see whether the prenatal approach still gives unbiased estimates under MAR.","tokens_in":42180,"feed_emoji":"🤰","tokens_out":6508,"duration_ms":65800,"temperature":0.7,"pith_summary":"This paper asks whether the way researchers find pregnancies in insurance claims and electronic health records—by looking for a recorded outcome like a delivery, versus by looking for a first prenatal visit—changes what they learn about prenatal medication effects. Using 10 million simulated pregnancies, it claims that when missing pregnancy outcomes are caused by unobserved miscarriage (missing not at random), the two strategies produce similarly biased risk differences and risk ratios. It claims further that the prenatal approach, which includes pregnancies whose outcomes are never observed, removes the bias only when the reasons outcomes are missing are measured variables that can be adjusted for. The practical message is that switching from outcome-based to prenatal pregnancy identification is not a cure for selection bias, but it is what makes sensible corrections possible for measurable missingness and makes the remaining bias quantifiable through sensitivity analyses.","feed_headline":"Prenatal drug studies get biased by how pregnancies are found","feed_subtitle":"Simulating 10 million pregnancies, only enrolling all pregnancies fixes bias, and only when missingness has measured causes.","key_machinery":"The machinery is the contrast between two cohort-construction designs: an outcome-based design that conditions on $S=1$ (observed outcome) and a prenatal design that enrolls all pregnancies at a first encounter and treats unobserved outcomes as missing data. Missingness is classified by the standard MCAR/MAR/MNAR trichotomy, and the analysis of observed pregnancies uses an Aalen-Johansen estimator within severity-rurality strata, treating miscarriage and non-preeclamptic live birth as competing events, then directly standardizes to the joint covariate distribution to estimate total effects. The load-bearing identity is that standardization can reweight the observed-pregnancies cohort only when the missingness mechanism is MAR; under MNAR, no reweighting recovers the target population, and a nonparametric bounds calculation over all possible outcomes of censored pregnancies is used to display the full range of possible effects.","core_discovery":"The central claim is that the outcome-based approach to pregnancy identification—building a cohort from observed deliveries or observed outcomes—conditions on a post-exposure event ($S=1$) and cannot generally be repaired by analysis, whereas the prenatal approach reframes the problem as missing outcome data that can sometimes be fixed. In simulations where all missingness came from unobserved miscarriage, the three analytic samples (observed deliveries, observed outcomes, and observed pregnancies) all overestimated absolute risks and returned comparably biased risk differences (RDs) and risk ratios (RRs); at 20% missingness, log-transformed RR bias ranged from -0.12 to 0.33 for observed deliveries, -0.11 to 0.32 for observed outcomes, and -0.11 to 0.32 for observed pregnancies. When all missingness was due to measured covariates (hypertension severity and rurality), only the observed-pregnancies sample using standardization was unbiased, with log-RR bias from -0.02 to 0.01, while delivery-restricted and outcome-restricted samples remained biased. The paper therefore concludes that including pregnancies with unobserved outcomes does not by itself prevent bias, but it converts an intractable selection problem into a missing-data problem for which weighting, bounds, and sensitivity analyses are available.","pith_inferences":["If the simulation's idealization is relaxed so that first prenatal visits occur at varying gestational ages rather than exactly 9 weeks, the prenatal approach could inherit immortal-time or selection problems that the current setup excludes; this is a testable extension of the same simulation.","The same missing-data logic should transfer to other pregnancy outcomes with outcome-dependent ascertainment, such as birth defects diagnosed only after live birth, although the bias magnitudes will differ.","In real claims data, the 'all pregnancies' cohort may itself be incomplete because some pregnancies end before any prenatal claim appears, so the advantage shown here likely represents an upper bound unless linkage or multiple data sources capture those early losses.","The framing suggests that published outcome-based pharmacoepidemiology studies could be systematically re-analyzed after re-identifying pregnancies prenatally, to check whether effect estimates move in the directions predicted here."],"forward_implications":["Studies of early-pregnancy exposures (before 13 weeks) should expect some bias whenever exposure changes miscarriage risk or shares causes with miscarriage; sensitivity analyses, not cohort restriction, are needed.","Restricting to observed deliveries, a common default, is not a remedy and in these simulations was often the most biased sample.","When missing outcomes are attributable to measured factors, a prenatal cohort plus inverse-probability weighting or standardization can recover unbiased total effects; this is unavailable to outcome-based cohorts because the missing pregnancies were never identified.","Including pregnancies with unobserved outcomes gives investigators the extent of missingness, enabling nonparametric bounds that contained the true RD and RR in all simulated scenarios.","If treatment affects neither miscarriage nor preeclampsia, all identification strategies return unbiased effect estimates in the scenarios studied."],"supporting_citations":[{"why":"Establishes that real-world healthcare data are used for prenatal drug safety studies and motivates the outcome-based pregnancy identification problem.","marker":"[1]"},{"why":"Documents the recent shift toward identifying pregnancies at first encounter, motivating the prenatal approach.","marker":"[2]"},{"why":"Demonstrates that restricting to live births biases pregnancy-drug effect estimates, the prior simulation this paper extends.","marker":"[7]"},{"why":"Supplies causal-diagram rules for missing data that ground the MAR/MNAR distinction.","marker":"[16]"},{"why":"Provides nonparametric bounds for the risk function used to bracket estimates for censored pregnancies.","marker":"[23]"},{"why":"Supplies the Aalen-Johansen competing-risks estimator used to analyze the observed-pregnancies sample.","marker":"[34]"}],"fun_headline_variants":["Pregnancy detection method biases prenatal exposure effects","How pregnancies are found skews prenatal exposure studies","Outcome-based pregnancy ID biases prenatal study results","All-pregnancy cohort fixes bias only if missingness measured"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing simplification is that every pregnancy has a first prenatal encounter at exactly 9 weeks, so the prenatal cohort captures the entire target population at a common time zero; if real-world first-visit timing varies or some pregnancies are never seen prenatally, the prenatal approach may no longer identify the full target population and the advantage shown here may not hold.","fun_headline_variants_meta":{"raw":{"variants":["Pregnancy detection method biases prenatal exposure effects","How pregnancies are found skews prenatal exposure studies","Outcome-based pregnancy ID biases prenatal study results","All-pregnancy cohort fixes bias only if missingness measured"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000928,"raw_usage":{"total_tokens":4081,"prompt_tokens":1157,"completion_tokens":2924,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":773,"completion_tokens_details":{"reasoning_tokens":2864}},"tokens_in":773,"tokens_out":2924,"duration_ms":23331,"temperature":1.0,"reasoning_tokens":2864,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:32:22.698635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Analyze a real claims or electronic health record database linked to a complete pregnancy registry that records miscarriages and deliveries for the same source population, estimate a prenatal drug effect three ways (observed deliveries, observed outcomes, observed pregnancies with standardization), and check whether the observed-pregnancies estimate becomes unbiased when all measured predictors of missingness are adjusted for; if it remains biased, the central claim is falsified. A cheaper check is analytic: modify the simulation so first prenatal visits occur at varying gestational ages rather than exactly 9 weeks and see whether the prenatal approach still gives unbiased estimates under MAR.","supporting_citations":[{"cited_title":"Use of real-world evidence from healthcare utilization data to evaluate drug safety during pregnancy","cited_arxiv_id":null,"evidence_quote":"Establishes that real-world healthcare data are used for prenatal drug safety studies and motivates the outcome-based pregnancy identification problem."},{"cited_title":"Identification of pregnancies in healthcare data: A changing landscape","cited_arxiv_id":null,"evidence_quote":"Documents the recent shift toward identifying pregnancies at first encounter, motivating the prenatal approach."},{"cited_title":"Bias from restricting to live births when estimating effects of prescription drug use on pregnancy complications: A simulation","cited_arxiv_id":null,"evidence_quote":"Demonstrates that restricting to live births biases pregnancy-drug effect estimates, the prior simulation this paper extends."},{"cited_title":"Using causal diagrams to guide analysis in missing data problems","cited_arxiv_id":null,"evidence_quote":"Supplies causal-diagram rules for missing data that ground the MAR/MNAR distinction."},{"cited_title":"Nonparametric Bounds for the Risk Function","cited_arxiv_id":null,"evidence_quote":"Provides nonparametric bounds for the risk function used to bracket estimates for censored pregnancies."},{"cited_title":"An Empirical Transition Matrix for Non-Homogeneous Markov Chains Based on Censored Observations","cited_arxiv_id":null,"evidence_quote":"Supplies the Aalen-Johansen competing-risks estimator used to analyze the observed-pregnancies sample."}],"review_version":1}