{"id":"59aa4429-082b-4c3e-980e-77b5961bfbf3","arxiv_id":"2412.13234","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An observed-trajectory microsimulation method using EHR data usually outperforms a multi-state-model method and a Markov cohort model for estimating cost-effectiveness of cancer therapy sequences.","lead":"This paper develops two microsimulation methods for estimating the cost-effectiveness of cancer therapy sequences from electronic health record data, and tests them on synthetic and real bladder-cancer datasets. One method uses model-based transition probabilities; the other uses observed patient trajectories, which the authors find is often more accurate.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic validation does not test informative censoring, so the trajectory method's reported coverage may not hold in real EHR data where censoring correlates with prognosis.","rationale":"The paper is a careful methods study with an honest synthetic-data validation, and the trajectory method is a plausible practical tool. The reader's CONDITIONAL verdict is appropriate. My stress-test focuses on the single assumption that, if false, most directly undermines the trajectory method's validity: non-informative censoring. The synthetic data deliberately avoids this by construction, so the reported coverage in Table 2 cannot be extrapolated to EHR data without further evidence. I agree with the reader's weakest_assumption. I do not see an internal inconsistency in the method; the concern is an untested external validity condition. The proposed test is feasible because the code is public and the synthetic-data framework already exists; adding an informative-censoring arm would directly quantify the sensitivity. If the test shows major coverage failure, the verdict should move toward REJECT or at least require the authors to add caveats; if coverage holds, the current CONDITIONAL verdict stands with a request for sensitivity analyses.","tokens_in":15437,"tokens_out":6792,"duration_ms":61062,"concrete_test":"Extend the synthetic-data experiment in Section 3 by generating informative censoring. For each patient, draw a latent frailty U (e.g., log-normal) that multiplies all transition hazards, and draw the censoring time from a distribution whose rate is an increasing function of U (or of the time-varying health state). Run the same trajectory and mstate microsimulations with the same 100 bootstrap replicates and compare coverage of ΔCost, ΔEffect, and NMB against the true values from the uncensored full-cohort datasets. If NMB coverage drops below the nominal 95% for the realistic mechanisms, the non-informative-censoring assumption is load-bearing and the Conclusions should be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The trajectory method (Section 2.2) uses each patient's observed transitions only until censoring, then switches to multi-state-model transition probabilities for the remainder of the simulation. This is valid only if censoring is non-informative given covariates and prior transitions. The synthetic evaluation (Section 3.1) draws censoring times from a uniform distribution independent of outcomes, so Table 2's coverage rates reflect a best case. In EHR data, censoring is often driven by clinical deterioration, hospice referral, or transfer of care, all of which are correlated with prognosis. Under informative censoring, the model-based tail is applied to a selected subgroup, and the IPTW weights do not correct for this; the mstate transition probabilities themselves are estimated from data subject to the same censoring, so the tail can be biased. The paper neither tests this mechanism nor reports sensitivity analyses. Thus the central claim that the trajectory approach 'mostly produced confidence intervals that covered known values' is not established for the realistic EHR setting.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops two microsimulation approaches for estimating costs, QALYs, and net monetary benefit of cancer therapy sequences from EHR data: a multi-state model (mstate) approach that uses fitted transition probabilities, and a trajectory approach that uses each patient's observed transitions until censoring and then switches to mstate transition probabilities. Both methods are evaluated in synthetic EHR-like datasets generated with copula-based within-patient dependence and known true outcomes, and are applied to a Flatiron Health bladder cancer cohort comparing cisplatin/gemcitabine versus carboplatin/gemcitabine followed by immunotherapy. The paper reports that the trajectory approach mostly produced confidence intervals covering true values, that the mstate approach was biased, and that both outperform a homogeneous Markov cohort model.","tokens_in":15481,"tokens_out":6262,"duration_ms":58649,"significance":"If the trajectory method is valid, it provides a practical template for using longitudinal EHR data in cost-effectiveness analyses of treatment sequences when randomized sequence trials are unavailable, with the important advantage of preserving within-patient dependence in observed transitions. The paper's strengths include publicly available code, a synthetic evaluation design with known truth, careful calibration against IPTW survival curves, and incorporation of external utilities, costs, and adverse-event rates. The copula-based simulation framework is a useful stress test for within-patient correlation. However, the central coverage claim is not yet established for realistic EHR settings because informative censoring is not tested, and the base-case Clayton-copula results show the trajectory method missing the true cost and effect differences.","major_comments":[{"comment":"In the base-case Clayton copula scenario, the trajectory method's 95% confidence intervals for the cost difference (-$18,208 to $6,988) and the effect difference (0.01; CI -0.08 to 0.11) exclude the true values ($10,503 and 0.15, respectively). Since Clayton is described in Section 3.1 as the base case for the synthetic evaluation, this is not a peripheral scenario, and it directly weakens the Abstract's claim that the trajectory approach 'mostly produced confidence intervals that covered known values.' Please provide a per-scenario coverage summary and either restrict the claim to the scenarios where coverage holds or explain why the method fails in the very setting (within-patient dependence) it was designed to capture.","section":"Section 3.2, Table 2 (Clayton row)"},{"comment":"The trajectory method's validity depends on the assumption that censoring is non-informative conditional on covariates and prior transitions, because after censoring the simulation switches to multi-state-model transition probabilities. The synthetic evaluation in Section 3.1 draws censoring times from a uniform distribution independent of outcomes, so the coverage results in Table 2 represent a best case. In EHR data, censoring can be driven by clinical deterioration, hospice referral, or transfer of care, which are correlated with prognosis. The paper neither tests this mechanism nor reports sensitivity analyses under outcome-dependent censoring. Please add a simulation scenario with censoring times dependent on prognosis (or on latent frailty) and show the trajectory method's coverage, or provide a formal theoretical argument that the IPTW weighting and the observed-trajectory construction account for informative censoring. Without this, the central claim that the trajectory approach is well-calibrated in real EHR data is not established.","section":"Section 2.2 and Section 3.1"},{"comment":"The Discussion states 'we did not formally assess coverage,' yet the Abstract and Section 3.2 claim that the trajectory approach 'mostly produced confidence intervals that covered known values.' Because only one synthetic dataset is generated per scenario, the assessment is based on a single interval per scenario rather than a repeated-sampling coverage rate. Please clarify the distinction and, if a coverage claim is intended, report a formal coverage analysis based on repeated synthetic datasets; otherwise temper the language in the Abstract and Results to reflect the informal nature of the comparison.","section":"Section 5 and Section 3.2"},{"comment":"In several plausible scenarios, both microsimulation methods produce cost and effect differences with the wrong sign: for example, in the small-sample (Clayton) scenario, the trajectory method estimates ΔEffect = -0.07 (CI -0.24 to 0.10) against a true value of 0.16, and in the log-normal (Clayton) scenario it estimates ΔCost = -$11,573 (CI -$26,148 to $3,003) against a true value of $4,250. The Discussion acknowledges this bias, but the Conclusions state that both methods 'offer superior performance to a homogeneous Markov cohort approach,' which is true only relative to a very poor comparator. Please qualify the recommendation by specifying the conditions under which the trajectory method is reliable (e.g., sample size, survival distribution shape, and copula structure), since the current framing overstates the method's general applicability.","section":"Section 3.2, Table 2 (small-sample, log-logistic, log-normal rows)"}],"minor_comments":[{"comment":"The multi-state Cox model equation contains the string '𝑒𝑒𝑒𝑒𝑒𝑒' where the exponential function is intended; please replace it with the correct mathematical notation for exp().","section":"Section 2.1, equation"},{"comment":"The text says 'We ﬁt Weibull models with 6 different copula structures: Independence, Clayton, Gaussian compound symmetric, Gaussian unstructured, T compound symmetric, T unstructured.' Independence is not a copula structure; please reword to distinguish the independence case from the five copula models.","section":"Section 3.1, second paragraph"},{"comment":"The description of the bootstrap for the trajectory approach says patients are resampled within each treatment arm and the microsimulation is run with varying seeds, but it does not state whether the propensity-score model is refit in each bootstrap replicate. Please clarify; if the IPTW weights are not recomputed, the bootstrap will understate uncertainty by ignoring estimation error in the weights.","section":"Section 2.2, trajectory bootstrap"},{"comment":"Table 3 is dense and difficult to parse, especially the columns comparing the full cohort with the treatment-sequence subgroup; consider reformatting or splitting the table for readability.","section":"Section 4, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically interesting and the synthetic validation is a worthwhile design, but the main coverage claim is currently too strong relative to the evidence: the base-case Clayton scenario fails for the trajectory method, and informative censoring is entirely untested. These are load-bearing issues that can be addressed with additional simulation scenarios and a more careful statement of the conditions under which the method is reliable. I do not see a need for rejection, but the revision must grapple with the censoring mechanism and the mixed table results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful methods paper for health economists who work with EHR data. The observed-trajectory microsimulation with IPTW and bootstrap is new as far as I can tell, and the synthetic evaluation with known truth is the right way to test such a tool. It deserves referee time.\n\nThe strongest part is the validation design. Generating synthetic EHRs with copula-driven within-patient dependence and comparing both microsimulation methods against known costs/effects/NMB is exactly what a methods paper needs. The result that the trajectory approach often covers truth while the mstate model misses even in the independence Weibull case is worth publishing in itself. The clinical bladder cancer example is a reasonable demonstration, not a definitive clinical finding, and the paper mostly treats it that way.\n\nThe soft spots are real, but proportionate. First, the synthetic censoring is uniform and independent of outcomes, while EHR censoring is often tied to deterioration, hospice, or transfer of care. The trajectory method's post-censoring tail is taken from the mstate model, so if censoring is informative given covariates and prior transitions, the tail is applied to a selected subgroup and the coverage numbers in Table 2 are a best case. The paper flags no test or sensitivity analysis on that point; an honest limitation but one that matters for the central claim.\n\nSecond, the mstate bias in the independence Weibull scenario is unexplained. That is a red flag for a seemingly standard multi-state application. The authors note calibration does not guarantee unbiased results, but they don't dig into why the Cox-based mstate fails when its own data-generating model is a proportional hazards Weibull. It might be the discrete-time cycle approximation or the exponential simplification mentioned in Section 2.1, but it's left unresolved.\n\nThird, the clinical example is demonstrative: no sensitivity analyses, one composite AE, external utilities and costs used without formal uncertainty. The authors concede this in the discussion. That's fine for a methods demonstration, but it limits what the bladder result says clinically.\n\nOn the citation pattern: the authors cite their own earlier microsimulation work and the relevant multi-state/copula literature appropriately. Self-citation here is not a problem.\n\nWho is this for? Health economists and decision modelers who want to build EHR-informed sequence models. I'd cite it if I worked in that space. It is not a conceptual breakthrough, but it is a competent and honest step forward. Verdict: accept for peer review, with a request that the authors either test informative censoring or soften the claim about the trajectory method's coverage.","headline":"A solid methodological contribution with an honest synthetic validation, but the headline coverage result rests on independent censoring that real EHR data will often violate.","tokens_in":16124,"tokens_out":2384,"would_cite":true,"duration_ms":22010,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replaying patients' recorded EHR transitions until censoring, then completing the tail with a multi-state model, estimates cancer therapy sequence costs and QALYs more accurately than standard Markov cohort models.","keywords":["microsimulation","therapy sequence","cost-effectiveness analysis","multi-state models","electronic health records","inverse probability treatment weighting","advanced cancer","quality-adjusted life years"],"falsifier":"Create a simulated dataset where patients are censored exactly when they get sicker or enter hospice, then check whether the trajectory method's intervals still cover the true costs and quality-adjusted survival; if they miss often, the method depends on censoring being unrelated to prognosis.","tokens_in":15097,"feed_emoji":"🩺","tokens_out":10937,"duration_ms":92644,"temperature":0.7,"pith_summary":"The paper tries to establish that electronic health records contain enough longitudinal information to support cost-effectiveness modeling of multi-line cancer therapy sequences, even though EHRs lack costs and quality-of-life measures. It proposes two microsimulation strategies: one that generates transitions from fitted multi-state survival models, and one that replays each patient's observed transition times until censoring and then switches to model-based probabilities. In synthetic datasets built to mimic EHR structure with correlated within-patient outcomes, the observed-trajectory version usually produced confidence intervals that covered the true cost and QALY differences, while the multi-state version was often biased. The authors conclude that both microsimulation methods outperform a homogeneous Markov cohort approach and that patient-level EHR data should be considered for therapy-sequence cost-effectiveness studies. A bladder-cancer demonstration yields a positive net monetary benefit for cisplatin-based first-line therapy followed by immunotherapy at $100,000 per quality-adjusted life year.","feed_headline":"Real patient trajectories beat Markov models for treatment sequencing","feed_subtitle":"Replaying each patient's recorded treatment transitions gives more accurate cost and quality-adjusted survival estimates.","key_machinery":"The central object is a discrete-time health-state transition model with states Line 1, Line 2, Extensive Disease, and Death, run as a patient-level microsimulation. Two engines drive it: an mstate engine using Cox proportional-hazards multi-state models, which estimate hazards for each possible transition between health states, to compute subject- and time-specific transition probabilities, and a trajectory engine that replays each patient's observed EHR transition times until censoring and then uses the multi-state model for the tail. Propensity-score inverse probability of treatment weighting with average treatment effect weights rebalances the treatment arms, and bootstrap resampling of the cohort with different random seeds supplies standard errors that include both population and simulation variability.","core_discovery":"The central claim is that observed EHR trajectories carry more information about therapy sequence outcomes than fitted transition probabilities alone. The trajectory method forces each simulated patient to follow their recorded sequence of health states until censoring, setting the probability of the observed transition to 1 and all others to 0, then completes the remainder of follow-up with transition probabilities from the multi-state model. Inverse probability of treatment weighting with average treatment effect weights balances the non-randomized treatment arms, and bootstrap resampling of the cohort with different random seeds supplies standard errors that include both population and simulation variability. In nine synthetic scenarios with copula-induced within-patient correlation, the trajectory method mostly produced confidence intervals that covered the known true costs, quality-adjusted life years, and net monetary benefit, while the multi-state method frequently did not; both methods substantially outperformed a homogeneous Markov cohort model. The paper concludes that patient-level EHR-based data should inform cost-effectiveness models of therapy sequence when available.","pith_inferences":["If extended beyond the paper, the trajectory method's reliance on non-informative censoring means real-world applications should test sensitivity to informative censoring, such as patients leaving the EHR when they enter hospice or transfer care.","A neighbouring problem the method could address is comparative effectiveness of treatment sequences in other chronic diseases with ordered lines of therapy recorded in EHRs, such as heart failure or multiple sclerosis.","A testable extension would be to swap the multi-state model tail for cause-specific hazards or flexible parametric models; because the mstate approach alone was biased even under proportional hazards, the tail's specification is the likeliest source of residual bias."],"forward_implications":["Health economists can estimate the cost-effectiveness of alternative therapy sequences without randomized sequence trials, using EHR longitudinal data plus external cost and utility inputs.","The observed-trajectory method should be preferred over the multi-state model method for sequence questions when EHR progression data are available, because it was more accurate in the synthetic evaluation.","Homogeneous Markov cohort models are likely to underestimate costs and effects for sequence questions; the paper's synthetic results show both microsimulation methods improve on them.","The bladder cancer demonstration implies that cisplatin/gemcitabine followed by immunotherapy is cost-effective at $100,000 per QALY compared with carboplatin/gemcitabine followed by immunotherapy, under both modeling approaches.","Researchers can attach external cost, utility, and adverse-event values to health states, so the method works even though EHRs do not record those measures."],"supporting_citations":[{"why":"Provides the multi-state modeling framework used to estimate transition probabilities for the mstate approach and for the post-censoring tail of the trajectory approach.","marker":"[18]"},{"why":"Supplies the R implementation of multi-state model estimation and prediction used throughout the simulations.","marker":"[19]"},{"why":"Provides the inverse probability of treatment weighting method used to balance non-randomized treatment sequences.","marker":"[20]"},{"why":"Justifies the average treatment effect weighting that makes the IPTW survival curve directly comparable to microsimulation output.","marker":"[21]"},{"why":"Supplies the copula-based simulation procedure used to generate synthetic EHR-like data with within-patient correlated transitions.","marker":"[23]"},{"why":"Describes the EHR-derived oncology database from which the bladder cancer cohort and covariate distributions were drawn.","marker":"[16]"},{"why":"Documents the real-world progression variable, abstracted from clinical notes, that defines the observed transitions used by the trajectory method.","marker":"[26]"},{"why":"Provides the homogeneous Markov cohort model that serves as the comparison baseline in the synthetic evaluation.","marker":"[25]"}],"fun_headline_variants":["Patient trajectories beat Markov models for cancer sequencing","EHR trajectories sharpen cost-effectiveness of cancer therapy sequences","Real patient paths improve microsimulation for therapy sequences","Trajectory-based microsimulation beats Markov for cancer cost-effectiveness","Observed trajectories beat fitted models for cancer therapy sequencing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that patients who stop being followed in the EHR would have had the same future transitions as similar patients who stay, so the model's after-censoring tail is not biased by the reasons they left.","fun_headline_variants_meta":{"raw":{"variants":["Patient trajectories beat Markov models for cancer sequencing","EHR trajectories sharpen cost-effectiveness of cancer therapy sequences","Real patient paths improve microsimulation for therapy sequences","Trajectory-based microsimulation beats Markov for cancer cost-effectiveness","Observed trajectories beat fitted models for cancer therapy sequencing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000584,"raw_usage":{"total_tokens":2771,"prompt_tokens":996,"completion_tokens":1775,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1698}},"tokens_in":612,"tokens_out":1775,"duration_ms":12915,"temperature":1.0,"reasoning_tokens":1698,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:28:44.194136+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Create a simulated dataset where patients are censored exactly when they get sicker or enter hospice, then check whether the trajectory method's intervals still cover the true costs and quality-adjusted survival; if they miss often, the method depends on censoring being unrelated to prognosis.","supporting_citations":[{"cited_title":"Differences in target estimands between different propensity score- based weights","cited_arxiv_id":null,"evidence_quote":"Justifies the average treatment effect weighting that makes the IPTW survival curve directly comparable to microsimulation output."},{"cited_title":"A simulation procedure based on copulas to generate clustered multi-state survival data","cited_arxiv_id":null,"evidence_quote":"Supplies the copula-based simulation procedure used to generate synthetic EHR-like data with within-patient correlated transitions."},{"cited_title":"Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US: Flatiron Health, SEER, and NPCR","cited_arxiv_id":null,"evidence_quote":"Describes the EHR-derived oncology database from which the bladder cancer cohort and covariate distributions were drawn."},{"cited_title":"Analysis of a Real-World Progression Variable and Related Endpoints for Patients with Five Different Cancer Types","cited_arxiv_id":null,"evidence_quote":"Documents the real-world progression variable, abstracted from clinical notes, that defines the observed transitions used by the trajectory method."},{"cited_title":"Markov Models and Cost Effectiveness Analysis: Applications in Medical Research","cited_arxiv_id":null,"evidence_quote":"Provides the homogeneous Markov cohort model that serves as the comparison baseline in the synthetic evaluation."}],"review_version":1}