{"id":"0bf8180b-d285-4d28-8937-75c347410f82","arxiv_id":"2505.01785","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"TV-SurvCaus combines recurrent sequence encoding, MMD representation balancing, and stabilized inverse probability weighting to estimate counterfactual survival curves under dynamic treatment regimes.","lead":"This paper introduces a deep learning method that estimates the effect of time-varying treatment plans on survival outcomes by learning balanced patient representations. It could matter for clinical research because it promises more accurate individualized treatment effect estimates in longitudinal data, though the current preprint's results are unverifiable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.6's consistency claim turns the needed convergence into Condition (iv): no argument shows the Eq. (10) minimizer drives dH(P^φ_a, P^φ_{a'}) to zero, so TV-PEHE→0 is unproven.","rationale":"The reader identified exactly the same weak point: Condition (iv) of Theorem 4.6 presupposes the asymptotic balancing success that TV-PEHE→0 is supposed to prove. My reading of the manuscript confirms this is the most load-bearing gap. If condition (iv) were a theorem rather than an assumption, the consistency argument would still need the factual weighted loss to identify F*_Ti and a transfer step from balance to counterfactual risk; but as it stands, the theorem's stated assumptions contain the conclusion. The empirical placeholder sentence in Section 7 and the non-existent synthetic-data description further mean the experimental claims cannot be checked, so rejection of the current version is warranted. I do not see a need to change the reader's REJECT verdict; the paper may become salvageable with a real proof, released code, and completed experiments.","tokens_in":15865,"tokens_out":6003,"duration_ms":67637,"concrete_test":"Construct the simplest K=1 feedback DGP satisfying Assumptions 3.1–3.5 (baseline W, A0, X1=βA0+ε, A1|X1, survival under each regime). Train the actual objective in Eq. (10) with Gaussian MMD for a sequence of α values (including α→∞) and n growing; compute dH(P^φ_n_a, P^φ_n_{a'}) and TV-PEHE. If dH stays bounded away from zero for any finite α, or if TV-PEHE does not vanish even when dH is driven down, then Condition (iv) is not a consequence of the proposed learning rule and Theorem 4.6 has no demonstration. Independently: re-derive Theorem 4.6 from the objective without quoting Condition (iv); the derivation must produce an explicit α_n and capacity condition under which the IPM converges.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Eq. (3) holds is supported only by Theorem 4.6, whose Condition (iv) is exactly that the balancing penalty drives the IPM discrepancy dH(P^φ_a, P^φ_{a'}) to zero as n grows. The proof sketch then 'combines effects' by citing that condition, so the theorem is conditional on the core mechanism it is meant to establish. It does not state how α in objective (10) must scale, what model capacity suffices, or how optimization error is overcome, and it assumes in Condition (iii) that stabilized weights are correct/consistently estimated and bounded—also a substantive asymptotic requirement. The paper also marks the experiments as placeholders ('(Assuming tables contain actual results now)' at the start of Section 7), describes the synthetic generator as 'described previously' without describing it, and labels the proof of Theorem 4.8 a 'conceptual decomposition.' Consequently, neither the theoretical nor the empirical pillar currently provides independent evidence for the claimed consistency and superiority.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TV-SurvCaus, a neural architecture for estimating the causal effect of time-varying treatments on survival outcomes. The method combines an RNN sequence encoder, a representation network trained with an MMD-based balancing penalty, and a discrete-time survival prediction head, with stabilized inverse probability weights in the loss. The theoretical sections claim a generalization bound connecting TV-PEHE to representation imbalance, variance analysis for sequential stabilized weights, consistency of the estimator, convergence rates under temporal dependence, and a bias decomposition for treatment-confounder feedback. The empirical section reports results on synthetic, semi-synthetic (SUPPORT, TCGA, METABRIC), and MIMIC-III data, claiming consistent improvements over Cox-MSM, G-formula, DeepSurv-L, RSF-MSM, and RMSN.","tokens_in":16064,"tokens_out":3868,"duration_ms":39075,"significance":"If the claims were established, the paper would offer a useful extension of representation balancing to dynamic treatment regimes with censored survival data, a setting that is indeed under-explored. The conceptual framing is relevant and the proposed architecture is plausible. However, the current manuscript provides no verifiable theoretical proof: the central consistency theorem is conditioned on essentially the conclusion, the convergence rate is an unspecified decomposition, and the bias theorem is explicitly labeled a conceptual decomposition. The empirical pillar is also missing: Section 7 opens with the placeholder text '(Assuming tables contain actual results now)', the synthetic data generator is never described, and no training details, hyperparameters, code, or data-processing steps are given. Consequently, the claimed significance cannot be assessed beyond the proposal level, and the paper in its present form does not support its central assertions.","major_comments":[{"comment":"The consistency claim TV-PEHE → 0 is conditioned on the assumption that the balancing penalty drives the IPM discrepancy d_H(P^φ_a, P^φ_a') to zero as n → ∞. This is essentially the conclusion the theorem purports to prove. The proof sketch does not show that the minimizer of the empirical objective in Eq. (10) achieves this; it merely cites Condition (iv) in the combining step. No argument is given for how α must scale, what model capacity suffices, or how optimization error is overcome. As stated, the theorem is circular and does not establish consistency of the TV-SurvCaus estimator.","section":"Section 4.3, Theorem 4.6, Condition (iv)"},{"comment":"The empirical section is not a real evaluation: the text explicitly says '(Assuming tables contain actual results now)'. The tables appear to be placeholders and cannot be independently checked. In addition, the synthetic data generator is described as 'described previously' but is never actually described anywhere in the manuscript, and no details are provided for the semi-synthetic constructions, the MIMIC-III cohort, hyperparameter selection, or experimental protocol. The central empirical claim that TV-SurvCaus outperforms baselines is therefore unsubstantiated in this submission.","section":"Section 7 (opening sentence and Tables 1-11)"},{"comment":"The claimed convergence rate is not derived. Equation (4) is a generic error decomposition into statistical error, approximation error, and balancing error, where the statistical rate R_stat(n, F, β) is left unspecified. The proof sketch does not prove any of the three displayed terms, and no explicit rate is given for the time-varying survival setting. Since the paper lists 'convergence rates for representation learning with temporal dependencies' as a contribution, this is a load-bearing gap rather than a minor omission.","section":"Section 4.4, Theorem 4.7"},{"comment":"The proof sketch is described as 'primarily a conceptual decomposition'. This is not a theorem with a proof; it is a qualitative list of bias sources. The statement does not define the bias components formally, does not state assumptions under which the decomposition holds, and provides no bound on the total bias. As such, Theorem 4.8 does not provide the 'refined analysis of bias' promised in the introduction.","section":"Section 4.5, Theorem 4.8"},{"comment":"The main theoretical motivation for the balancing objective is not established. Theorem 4.1 is stated as an adaptation of existing domain-adaptation bounds, but no proof is given that the IPM discrepancy between representation distributions across treatment sequences controls counterfactual survival risk under the weighted discrete-time survival loss with censoring. Corollary 4.2 handwaves the step 'relating risk to MSE (e.g., via properties of the survival loss)', which is precisely the step that would connect the abstract bound to the TV-PEHE metric. Without this step, the bound does not directly motivate the objective in Eq. (10).","section":"Section 4.1, Theorem 4.1 and Corollary 4.2"}],"minor_comments":[{"comment":"The notation T_i(−1) = empty is used in the setup, but Equation (2) uses T_i(k−1) for k = 0; please make the convention for the empty history explicit at the point where the weights are defined.","section":"Section 3.1"},{"comment":"The notation 'TV-CATES(X, a, a′, τ; ·) − TV-CATES(X, a, a′, τ; ·∗)' is confusing because the first argument uses the estimator and the second uses the true function; clarify by writing F̂ and F* explicitly.","section":"Section 3.3, Definition 3.4"},{"comment":"Theorem 4.3 is a definition of stabilized weights, not a theorem; it should be labeled as a definition or equation rather than a numbered theorem.","section":"Section 4.2, Theorem 4.3"},{"comment":"The synthetic data section says the data is 'Generated as described previously', but no previous description exists in the paper; the full generator including the treatment assignment mechanism, covariate dynamics, outcome generation, censoring, and feedback strength β must be specified.","section":"Section 6.1"},{"comment":"The table reports '28-day mortality risk reduction', but the method section does not explain how this estimand is computed from the survival curves; state the formula used.","section":"Section 7.3, Table 8"},{"comment":"The limitations section appropriately acknowledges assumptions and computational cost, but the paper should also state explicitly that no code or data are released, which is relevant for reproducibility.","section":"Section 8.3"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an incomplete draft rather than a finished submission. The explicit placeholder text in Section 7 ('(Assuming tables contain actual results now)') and the missing synthetic data description are not presentation issues; they indicate that the empirical portion of the paper does not exist yet. The theoretical contribution is likewise not in a publishable state, since the main consistency theorem is conditional on the convergence it is supposed to prove. Even under a generous reading, the central claims are not supported by the current text, so rejection, rather than major revision, is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: not a viable paper in this form. Section 7 opens with \"(Assuming tables contain actual results now)\", and the main consistency theorem assumes the convergence it is supposed to establish. This is a desk reject, not a paper to referee.\n\nWhat is genuinely useful: the problem is real, and the proposed architecture—an RNN history encoder, an MMD balancing term over treatment sequences, and a stabilized-IPTW-weighted survival likelihood—is a sensible combination of known building blocks. The paper engages fairly with the literature and states its own limitations (sequential exchangeability, positivity, hyperparameter sensitivity) without flinching. If someone actually built this and validated it, it could become a solid empirical paper.\n\nThe soft spots are structural. Theorem 4.6 says TV-PEHE goes to zero, but Condition (iv) is exactly the assumption that the balancing penalty drives the representation discrepancy to zero. That is the result, not a premise. The proof sketch just cites that condition as \"combining effects.\" Theorem 4.7 is not a rate; it is a generic bias-variance-balance decomposition with no explicit rates, and Theorem 4.8 is admittedly \"primarily a conceptual decomposition.\" So the theoretical contribution reduces to a restatement of Shalit/Johansson/Ben-David with a survival loss, which is fine as background but not as new theory. The synthetic data generator is \"described previously\" without being described. No code, no data, no real results anywhere. The tables contain numbers, but the paper itself tells the reader they are placeholders.\n\nWho gets value? Possibly someone scouting research directions, but even then the value is marginal. The architecture could be re-implemented, but nothing here makes the reported numbers trustworthy. This should be desk-rejected and sent back to the author with a clear list: replace the placeholders, prove the consistency claim without assuming it, describe the data generating process, and release code. Then it might be worth another look.","headline":"A plausible architecture with honest limitations, but no actual results: the experiments are explicit placeholders and the main consistency theorem assumes what it claims to prove.","tokens_in":16623,"tokens_out":2964,"would_cite":false,"duration_ms":30188,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62N01","62N02","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"TV-SurvCaus claims that balancing sequence-level representations across treatment histories makes counterfactual survival functions consistently estimable, with TV-PEHE converging to zero under stabilized weighting and an MMD regularizer.","keywords":["time-varying treatments","causal survival analysis","representation balancing","treatment-confounder feedback","potential outcomes","maximum mean discrepancy","TV-PEHE","marginal structural models"],"falsifier":"On synthetic data with a known true survival function, induce strong time-varying confounding and misspecify the propensity model used to build the stabilized weights; if the MMD between learned representations stays bounded away from zero and TV-PEHE does not shrink with increasing sample size, the consistency claim is refuted.","tokens_in":15610,"feed_emoji":"🩺","tokens_out":8107,"duration_ms":76269,"temperature":0.7,"pith_summary":"TV-SurvCaus sets out to show that representation balancing—learning a latent history representation that looks similar across different treatment sequences—extends to causal survival analysis with time-varying treatments and censored outcomes. The paper argues that estimating counterfactual survival curves under dynamic treatment regimes is feasible and provably consistent when a survival loss weighted by stabilized inverse-probability weights is combined with an MMD balancing penalty. It backs this with a generalization bound on the time-varying precision in estimating heterogeneous effects (TV-PEHE), a consistency theorem, convergence-rate discussion, and a bias decomposition. If correct, the framework would give medical and policy analysts a principled way to estimate individual treatment effects of evolving treatment protocols from longitudinal observational data, including real clinical records.","feed_headline":"Time-varying survival effects become consistently estimable","feed_subtitle":"Representation balancing across treatment sequences plus stabilized weighting closes the counterfactual error gap.","key_machinery":"The load-bearing object is the balanced representation $z = \\phi(\\psi(X,T))$, built by a sequence encoder (an LSTM or GRU) followed by a representation network, with maximum mean discrepancy (MMD) as the integral probability metric measuring distributional balance between treatment sequences. The training objective combines a discrete-time survival negative log-likelihood weighted by stabilized sequential inverse-probability weights, the MMD balancing penalty $\\alpha L_{bal}$, and an $L_2$ regularization term. The main theoretical tool is a domain-adaptation-style bound on counterfactual risk; the bound identifies representation imbalance as the term the balancing loss must drive to zero, and the consistency theorem rests on that driving condition.","core_discovery":"The central discovery is that the counterfactual survival risk under an alternative treatment sequence is bounded by the factual risk plus an integral-probability-metric discrepancy between representation distributions; minimizing that discrepancy and using bounded stabilized inverse-probability weights makes the TV-PEHE converge in probability to zero (Theorem 4.6). In the paper's own terms, the estimated potential survival function converges to the true function for sequences represented in the data, and experiments on synthetic, semi-synthetic, and real clinical data show lower TV-PEHE and better discrimination and calibration than marginal structural models, G-formula, and deep sequence baselines.","pith_inferences":["In finite samples the practical value probably hinges on the balancing strength $\\alpha$; a testable prediction is that as $\\alpha$ shrinks to zero, TV-SurvCaus degrades toward a plain sequence survival model, and monitoring validation MMD should track residual confounding bias.","The same bounding structure suggests extensions to continuous treatment doses and competing-risk endpoints, provided the IPM and survival head are redefined; the representation-balancing argument itself would carry over.","The theory's reliance on asymptotic balance implies that overlap violations are the real threat in applications: when some histories make treatment nearly deterministic, neither weighting nor representation balancing can recover the counterfactual, and reported causal effects in such strata should be downweighted or omitted."],"forward_implications":["Counterfactual survival curves for arbitrary treatment sequences can be estimated from observational longitudinal data without fitting a separate outcome model per regime, since one balanced representation plus a sequence-conditioned prediction head serves all sequences.","In nonlinear data-generating processes and longer treatment horizons, the combination of stabilized weighting and representation balancing gives progressively larger reductions in TV-PEHE relative to marginal structural models and G-formula.","The balancing regularizer absorbs part of the bias from treatment-confounder feedback, so the estimator can remain competitive when the propensity model used for weighting is mildly misspecified.","Consistent TV-PEHE estimates make downstream dynamic treatment regime optimization possible, because the learned counterfactual survival models supply the per-sequence outcomes such methods require."],"supporting_citations":[{"why":"Supplies the domain-adaptation generalization bound that Theorem 4.1 adapts to counterfactual survival risk.","marker":"Ben-David et al. [2010]"},{"why":"Introduced representation balancing for static individual treatment effect estimation, the approach extended here to time-varying treatments.","marker":"Shalit et al. [2017]"},{"why":"Provides the generalization-bound and M-estimator machinery used to connect representation discrepancy to the TV-PEHE objective.","marker":"Johansson et al. [2020]"},{"why":"Defines marginal structural models and inverse-probability weighting, which motivate the stabilized survival loss.","marker":"Robins et al. [2000]"},{"why":"Recurrent marginal structural networks are the closest deep baseline that combines RNNs with IPTW, against which TV-SurvCaus is evaluated.","marker":"Lim et al. [2018]"},{"why":"Provides the discrete-time nnet-survival architecture used as the survival prediction head.","marker":"Gensheimer and Narasimhan [2019]"},{"why":"Supplies the real-world critical-care database used in the empirical evaluation.","marker":"Johnson et al. [2016]"},{"why":"Previous SurvCaus work that introduced representation balancing for static survival causal inference and is extended to the time-varying setting.","marker":"Abraich et al. [2022]"}],"fun_headline_variants":["TV-SurvCaus: balanced representations for survival causal inference","Time-varying treatments: causal survival estimates with less bias","Dynamic balancing makes survival counterfactuals converge","Representation balancing improves dynamic survival effect estimation","Sequential balancing stabilizes causal survival analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The consistency proof requires that the balancing penalty can actually drive the distributional discrepancy between treatment groups to zero as the sample grows, and that the stabilized inverse-probability weights are correctly specified or consistently estimated and bounded; if either condition fails, the claimed convergence of the estimator and of TV-PEHE to zero does not follow.","fun_headline_variants_meta":{"raw":{"variants":["TV-SurvCaus: balanced representations for survival causal inference","Time-varying treatments: causal survival estimates with less bias","Dynamic balancing makes survival counterfactuals converge","Representation balancing improves dynamic survival effect estimation","Sequential balancing stabilizes causal survival analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1290,"prompt_tokens":870,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":347}},"tokens_in":486,"tokens_out":420,"duration_ms":4642,"temperature":1.0,"reasoning_tokens":347,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:10:05.159233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On synthetic data with a known true survival function, induce strong time-varying confounding and misspecify the propensity model used to build the stabilized weights; if the MMD between learned representations stays bounded away from zero and TV-PEHE does not shrink with increasing sample size, the consistency claim is refuted.","supporting_citations":[],"review_version":1}