{"id":"bfc9953a-bc1f-4495-93ae-8ae113cfbef7","arxiv_id":"2507.06961","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A re-weighted inverse probability value estimator for off-policy evaluation under non-ignorable missing data, with consistency and normality guarantees.","lead":"Off-policy evaluation estimates how well a new treatment policy would perform using data collected under older policies. This paper shows that when some outcomes are missing, standard estimates stay valid only if missingness is random, and it proposes a re-weighted estimator with confidence intervals for the harder non-random case.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed T→∞ consistency/asymptotic normality cannot hold under monotone dropout: positivity (A.2(b)) makes observed trajectory lengths finite, so nT is not the effective sample size.","rationale":"The paper's core finite-horizon story—CC is biased under MNAR, IPW with a correctly specified dropout propensity and valid shadow variable fixes the bias as n grows—is plausible and well supported by simulations. The proofs are detailed, and the n→∞ direction appears consistent with the estimating-equation arguments. However, the theorems explicitly advertise a bidirectional asymptotic theory, and that T→∞ direction is not merely unproven; it is inconsistent with the monotone missingness model the paper itself defines. With per-step observation probability bounded below by cλ (A.2(b)), each trajectory's observed length is finite with bounded mean, so increasing the nominal horizon T adds no new information for fixed n. The Appendix G proofs import nT as the effective sample size from Shi et al. (2021b), but dropout censoring breaks that identification. The reader's weakest assumption—correct specification of the dropout propensity and validity of the shadow variable—is a real and acknowledged limitation, but the asymptotic-regime gap is more decisive because it makes a stated theorem false as written rather than merely fragile under misspecification. The paper should be accepted only after the theorems are corrected (e.g., n→∞ only, or a joint regime with dropout rate shrinking) or the bidirectional claim is removed. No ad hominem intent; this is a technical overclaim in an otherwise substantive framework.","tokens_in":47054,"tokens_out":8038,"duration_ms":99196,"concrete_test":"Run the Section 5 MNAR setting with n fixed (e.g., 500) and T = 25, 50, 100, 200, 400 using the paper's IPW(P) estimator over 250 replications. Record average bias, MSE, CI width, and average observed trajectory length. If bias and CI width do not shrink as T increases while nominal nT grows, and mean observed length stabilizes at a constant independent of T, the claimed T→∞ consistency/asymptotic normality is refuted. Analytically, compute E[N_obs] ≤ n/cλ < ∞, showing the estimator cannot achieve a √(nT)-rate CLT in T alone.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorems 4.6 and 4.7 state consistency and √(nT)-asymptotic normality \"as either n→∞ or T→∞\". Under the paper's own monotone missingness setup, the T→∞ leg is not valid. For a fixed subject, dropout time C satisfies P(C=t | ...)=λ(...), and Assumption A.2(b) gives 1-λ ≥ cλ > 0, so C is stochastically dominated by a geometric random variable and has finite mean. Thus for fixed n, the number of actually observed transition tuples N_obs = Σ_i (C_i ∧ T) is bounded in probability as T→∞; it does not diverge. Consequently the IPW estimator is based on O_p(1) observations, cannot converge in probability to V^π merely as T→∞, and the √(nT) normalization in Theorem 4.7 is inappropriate. The same objection applies to the CC consistency statement in Theorem 4.5. The proofs in Appendix G repeatedly treat nT as the sample size in moment and CLT arguments inherited from Shi et al. (2021b), which is only legitimate when T is the length of complete trajectories. With dropout, the effective sample size is Σ η_{i,t+1}, not nT. This does not disprove consistency as n→∞ with T fixed, but it means the \"bidirectional\" theorem, a headline contribution, is overclaimed as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies off-policy evaluation (OPE) in infinite-horizon MDPs when logged trajectories are subject to monotone missingness. It argues that complete-case (CC) value estimation remains consistent under ignorable missingness (MAR) but can be biased under nonignorable missingness (MNAR). To restore consistency, the paper proposes an inverse probability weighted (IPW) value estimator, with the dropout propensity estimated through a shadow-variable estimating equation, and claims bidirectional consistency and asymptotic normality as either the number of trajectories n or the horizon T tends to infinity. Supporting results include simulations, a synthetic sepsis experiment, and an application to MIMIC-III data.","tokens_in":47317,"tokens_out":13522,"duration_ms":172313,"significance":"The paper addresses a genuine gap: OPE under MNAR is rarely studied, and the proposed IPW framework with shadow variables is a natural and potentially useful extension of the missing-data literature to sequential decision problems. The negative result that CC estimators can be biased under MNAR is valuable, and the MAR consistency argument is plausible under the stated conditional independence. The authors provide extensive simulations, a real-data application, and released code. However, the headline bidirectional theorems are not supported by the current proofs in the T→∞ direction under the formulated dropout model, and the dropout propensity estimating equation is not well defined for time points after dropout. These issues are load-bearing for the paper's central claims and require substantive correction.","major_comments":[{"comment":"The claim that consistency and √(nT)-asymptotic normality hold 'as either n→∞ or T→∞' is not supported by the proof under the paper's monotone dropout setup. The proofs in Appendix G treat nT as the number of terms in the moment sums (e.g., Eqs. (29)-(31) and the martingale construction in Step 3 of the proof of Theorem 4.7), but under monotone missingness the number of observed transitions is N_obs = Σ_i (C_i ∧ T), not nT. If dropout is eventually certain, as in the paper's simulation models where the dropout probability is bounded away from zero, N_obs = O_p(n) for fixed n and the estimator cannot converge as T→∞ at the claimed √(nT) rate; if dropout is not eventually certain, an additional condition and a different variance normalization are needed. The theorems should be restricted to n→∞ with T fixed, or reformulated with N_obs and an explicit assumption that the effective sample size diverges in the stated asymptotic regime.","section":"Section 4.2, Theorems 4.5-4.7; Appendix G"},{"comment":"The estimating equation for ψ is not well defined as written for subjects who have dropped out. For a subject with dropout time C_i and for all t > C_i, we have η_{i,t+1}=0 and hence m_i,t = -h(S_i,t, A_i,t, Z_i,t), which requires the state, action, and shadow variable at time t even though these are not observed (and are not generated if the trajectory terminates at dropout). The paper never restricts the sum in Eq. (4) to person-time with η_{i,t}=1, nor does it specify that h is set to zero after dropout. Consequently, the left-hand side of Eq. (4) cannot be computed from the observed data, and the identifiability of ψ through the shadow variable is not established for the actual observed data. This also affects the expansion of √(nT)(bψ - ψ*) in Appendix G.2. Please rewrite Eq. (4) over observed person-time and re-derive the subsequent theory under that definition.","section":"Section 4.3, Eq. (4); Appendix B"},{"comment":"For the semi-parametric dropout model, the implemented variance estimator replaces Ω_IPW by eΩ_IPW, which drops the H2 term that accounts for estimation of ψ. Theorem 4.7, however, is stated for bσ^2 in Eq. (34), and no result in the paper shows that the omitted term is asymptotically negligible. Since Algorithm 1 and all real-data and semi-parametric simulation results use the approximation (35), the coverage guarantee for the IPW(SP) estimator is not proven. Either prove that the approximation error is o_p(1) under the semi-parametric model, or state Theorem 4.7 only for the parametric dropout model with the full variance estimator and present eΩ_IPW as a heuristic approximation.","section":"Appendix G.3, Eqs. (33)-(35)"}],"minor_comments":[{"comment":"The column heading says 'standard error in parenthesis', but the numbers appear to be Monte Carlo standard deviations of the bias rather than standard errors of the mean; please clarify the definition and report standard errors of the Monte Carlo averages if that is the intent.","section":"Section 5, Table 1"},{"comment":"The x-axis label 'Confidence level 1' should read 'Nominal confidence level 1−α'.","section":"Figure 2"},{"comment":"The typeset equation contains the LaTeX underbrace annotation '| {z }' inside the matrix expression; this should be removed.","section":"Equation (3)"},{"comment":"The shadow variable is described in Section 6 as the previous GCS score, S^GCS_{t−1}, but Appendix E.3 writes Z_{t+1} = 1(S^GCS_{t+1} ≥ 14); please clarify the timing and explain whether the shadow variable is part of the state vector used in the value estimation.","section":"Section 6 and Appendix E.3"},{"comment":"The paper correctly acknowledges that when conditions for a chosen shadow variable are only partially satisfied, the observed likelihood is non-identifiable and the estimating equation is likely to fail; this important caveat should be connected more explicitly to Assumption A.2(c) and to the MIMIC-III application, where the validity of the chosen shadow variable is not empirically verified.","section":"Appendix B.3"}],"recommendation":"major_revision","confidential_remarks":"The main theoretical selling point is the bidirectional asymptotics, but that part of the argument is not currently sound: the proof's sample size is nT while the data contain only N_obs observed transitions after monotone dropout. The ill-defined estimating equation for ψ after dropout compounds the problem. I would recommend major revision rather than rejection because the n→∞ with fixed T direction is plausible and the missing-data-in-OPE problem is worth pursuing, but the authors need to substantially rework the asymptotic statements, the dropout estimation equation, and the variance estimator for the semi-parametric case."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the core result—that IPW restores consistency for OPE under nonignorable monotone dropout when the propensity is correctly specified and a valid shadow variable is available—is real and mostly well-proved for n→∞ with fixed horizon T. Second, the headline \"bidirectional\" theorems (4.6 and 4.7), which claim consistency and √(nT) normality as either n→∞ or T→∞, are overclaimed. Under Assumption A.2(b), dropout time is stochastically dominated by a geometric variable, so for fixed n the number of observed transitions stays bounded as T→∞; the effective sample size is Σ_i C_i, not nT. The proofs in Appendix G inherit Shi et al.'s nT-sample arguments, which are only legitimate when T is the length of complete trajectories. The T→∞ leg should be removed or replaced with a joint divergence condition where n and T grow together, with n dominating. This does not kill the n→∞, T-fixed consistency and CLT, which is what the simulations actually exercise.\n\nWhat is genuinely new: formalizing MAR vs MNAR for infinite-horizon OPE, proving that the complete-case estimator stays consistent under MAR, and giving an IPW estimator that accounts for dropout-propensity estimation in the variance through the H2 term in Eq. 31. That is a real step beyond single-stage missing-data work and beyond the prior OPE literature. The shadow-variable identification machinery is imported correctly. The paper is also honest about the shadow-variable verification problem; Appendix B.3 admits the estimating equation fails if the shadow-variable conditions are only partially satisfied.\n\nSoft spots, in order of severity. First, the bidirectional claim above; a referee should ask for a corrected theorem statement. Second, the semiparametric variance approximation in Eq. 35 drops the H2 term and is justified only empirically, and coverage does run a bit low at n=500. Minor-to-moderate, but it should be flagged. Third, the \"first\" claims are stronger than the cited Chu et al. 2023 supports; easy to soften. Fourth, the code is promised via a Github link that is not present; that needs fixing before publication.\n\nWho this is for: anyone doing OPE in healthcare or recommendation logs with informative dropout. The fixed-T identification and the MAR/MNAR distinction will be useful even after the asymptotics are tightened. Yes, this deserves serious peer review; the load-bearing idea survives the flawed T→∞ claim, and the proofs are detailed enough that a competent referee can isolate the issue. I would not desk-reject.","headline":"Real contribution for fixed-T OPE under MNAR, but the claimed T→∞ consistency and asymptotic normality are invalid under monotone dropout; needs theorem restatement and a linked codebase before I would sign off.","tokens_in":47890,"tokens_out":2669,"would_cite":true,"duration_ms":30426,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Nonignorable missingness makes ordinary off-policy value estimates biased, and an inverse-probability-weighted estimator built on a shadow variable restores consistency and valid confidence intervals.","keywords":["off-policy evaluation","nonignorable missingness","missing not at random","shadow variable","inverse probability weighting","dropout propensity","value function inference","reinforcement learning"],"falsifier":"Simulate MNAR data with a true dropout probability outside the fitted model class, for example a non-logistic function of the unobserved reward while fitting a logistic propensity, or generate the shadow variable so it also depends on dropout; under either misspecification, the IPW value estimate's bias should fail to vanish as $nT\\to\\infty$, contradicting Theorem 4.6.","tokens_in":46816,"feed_emoji":"🎯","tokens_out":15365,"duration_ms":139331,"temperature":0.7,"pith_summary":"Off-policy evaluation estimates the value of a target policy from logged trajectories, and logged trajectories are often cut short by dropout. The paper proves that when dropout is ignorable (missing at random), the usual complete-case estimator remains consistent; when dropout is nonignorable, because it depends on the unobserved reward or next state, the usual estimator is biased and the bias does not disappear as the sample grows. To fix that regime, the paper weights every observed transition by the inverse probability of observing it, estimating that probability through a moment equation that uses a shadow variable, an auxiliary covariate that predicts the unobserved outcome but is unrelated to dropout. It then proves the inverse-probability-weighted value estimator is consistent and asymptotically normal as either the number of trajectories or the horizon grows, and demonstrates in simulations and on a sepsis dataset that its confidence intervals attain nominal coverage.","feed_headline":"When dropout hides outcomes, inverse weighting saves policy evaluation","feed_subtitle":"Under informative missingness, a shadow-variable IPW estimator stays consistent with valid confidence intervals.","key_machinery":"The engine is the inverse-probability-weighted Bellman estimating equation $\\mathbb{E}[\\eta_{t+1}\\{1-\\lambda_{t+1}(\\psi)\\}^{-1} M_t(\\beta^\\pi)]=0$, where $M_t(\\beta^\\pi)=\\xi_t(R_{t+1}+\\gamma V^\\pi(S_{t+1})-Q^\\pi(S_t,A_t))$ is the Bellman residual and $\\xi_t$ is a linear sieve basis for the state-action features. The dropout propensity $\\lambda_{t+1}(\\psi)=\\lambda(S_t,A_t,R_{t+1},S_{t+1};\\psi)$ is identified through a shadow variable $Z_t$, a covariate that is conditionally independent of dropout given $(S_t,A_t,R_{t+1},S_{t+1})$ but associated with the unobserved outcome, and estimated from $\\mathbb{E}[(\\eta_{t+1}/\\{1-\\lambda_{t+1}(\\psi)\\}-1)h(S_t,A_t,Z_t)]=0$, in parametric or semiparametric exponential-tilting form. The load-bearing identity is $\\mathbb{E}[\\eta_{t+1}/\\{1-\\lambda_{t+1}(\\psi^\\ast)\\} \\mid \\mathcal{F}_t,R_{t+1},S_{t+1},\\eta_t=1]=1$, which restores the zero expectation of the Bellman residual at the true parameter and gives the estimator its consistency.","core_discovery":"The paper's central claim is that monotone nonignorable missingness breaks standard off-policy evaluation while ignorable missingness does not. Under Assumption A.1, the complete-case estimator $\\widehat V^\\pi_{\\mathrm{CC}}(G)$ is consistent under MAR (Theorem 4.5), but under MNAR it is biased because $\\mathbb{E}[\\eta_{t+1} M_t(\\beta^\\ast)] \\neq 0$. The proposed estimator weights each transition by $\\eta_{t+1}/\\{1-\\lambda(S_t,A_t,R_{t+1},S_{t+1};\\widehat\\psi)\\}$, with $\\widehat\\psi$ obtained from a shadow-variable moment equation; Theorem 4.6 gives consistency, and Theorem 4.7 gives $\\sqrt{nT}\\,\\widehat\\sigma^{-1}\\{\\widehat V^\\pi_{\\mathrm{IPW}}(G)-V^\\pi(G)\\} \\xrightarrow{d} N(0,1)$, where the limit holds as either $n\\to\\infty$ or $T\\to\\infty$. The paper also shows that the same inverse-weighting idea can be inserted into other OPE losses, such as fitted Q-evaluation.","pith_inferences":["Editorial inference: because the weighting identity only needs the conditional expectation of the response indicator, the same correction should carry over to fitted Q-evaluation and marginalized-importance-sampling estimators, but the paper only sketches those extensions and does not prove their asymptotics.","Editorial inference: the variance estimator used with a semi-parametric dropout model drops the uncertainty from estimating $\\psi$; users should expect mild under-coverage in small samples and could add a bootstrap or influence-function correction.","Editorial inference: when no credible shadow variable exists, the identification failure described in the paper implies that no MNAR-robust value estimate is possible without extra assumptions; a sensitivity analysis over candidate shadow variables is the practical counterpart."],"forward_implications":["Under missing-at-random dropout, complete-case value estimation needs no missingness correction: it remains consistent.","Under nonignorable dropout, complete-case estimates are biased and confidence intervals under-cover, with the bias persisting as $nT$ grows.","A correctly specified dropout propensity plus a valid shadow variable makes the IPW value estimator consistent and asymptotically normal, so uncertainty quantification is available.","The IPW correction is modular and can be combined with other value-estimation losses, including fitted Q-evaluation, rather than only the linear-sieve estimator in the main theorems."],"supporting_citations":[{"why":"Supplies the linear-sieve value-function inference whose Theorem 1 is the no-missing-data base that Theorems 4.5-4.7 extend.","marker":"Shi et al. (2021b)"},{"why":"Proves that the observed likelihood is non-identifiable under MNAR without an auxiliary variable, motivating the shadow-variable route.","marker":"Wang et al. (2014)"},{"why":"Establishes non-identifiability of the joint distribution when both the dropout propensity and outcome density are unknown.","marker":"Rotnitzky & Robins (1997)"},{"why":"Provides the semiparametric exponential-tilting model and estimating equations used to fit the dropout propensity.","marker":"Shao & Wang (2016)"},{"why":"Gives the non-parametric shadow-variable identification framework that justifies solving the dropout propensity from equation (4).","marker":"Miao et al. (2024)"},{"why":"Offers an alternative shadow-variable estimation procedure for nonignorable missing data.","marker":"Zhao & Ma (2022)"},{"why":"Defines the informative dropout propensity model that the paper adopts for the missingness mechanism.","marker":"Diggle & Kenward (1994)"},{"why":"Supplies the sepsis cohort, state/action/reward construction and clinical context for the real-data evaluation.","marker":"Komorowski et al. (2018)"},{"why":"Provides the MIMIC-III database from which the sepsis trajectories are extracted.","marker":"Johnson et al. (2016)"}],"fun_headline_variants":["Nonignorable missingness breaks OPE, inverse weighting fixes it","When missingness informs outcomes, IPW keeps OPE honest","Shadow variables rescue off-policy evaluation from informative dropout","Under informative missingness, IPW restores OPE consistency","Nonignorable missing data? Weight by shadow variables for valid OPE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The consistency proof collapses if the dropout model is not the true mechanism or if the chosen shadow variable is not conditionally independent of dropout given the unobserved outcome, because then equation (4) does not recover the dropout probabilities.","fun_headline_variants_meta":{"raw":{"variants":["Nonignorable missingness breaks OPE, inverse weighting fixes it","When missingness informs outcomes, IPW keeps OPE honest","Shadow variables rescue off-policy evaluation from informative dropout","Under informative missingness, IPW restores OPE consistency","Nonignorable missing data? Weight by shadow variables for valid OPE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2756,"prompt_tokens":910,"completion_tokens":1846,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1759}},"tokens_in":526,"tokens_out":1846,"duration_ms":12399,"temperature":1.0,"reasoning_tokens":1759,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:51:01.861257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate MNAR data with a true dropout probability outside the fitted model class, for example a non-logistic function of the unobserved reward while fitting a logistic propensity, or generate the shadow variable so it also depends on dropout; under either misspecification, the IPW value estimate's bias should fail to vanish as $nT\\to\\infty$, contradicting Theorem 4.6.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proves that the observed likelihood is non-identifiable under MNAR without an auxiliary variable, motivating the shadow-variable route."},{"cited_title":"and Robins, J","cited_arxiv_id":null,"evidence_quote":"Establishes non-identifiability of the joint distribution when both the dropout propensity and outcome density are unknown."},{"cited_title":"and Wang, L","cited_arxiv_id":null,"evidence_quote":"Provides the semiparametric exponential-tilting model and estimating equations used to fit the dropout propensity."},{"cited_title":"J., and Geng, Z","cited_arxiv_id":null,"evidence_quote":"Gives the non-parametric shadow-variable identification framework that justifies solving the dropout propensity from equation (4)."},{"cited_title":"and Kenward, M","cited_arxiv_id":null,"evidence_quote":"Defines the informative dropout propensity model that the paper adopts for the missingness mechanism."},{"cited_title":"A., Badawi, O., Gordon, A","cited_arxiv_id":null,"evidence_quote":"Supplies the sepsis cohort, state/action/reward construction and clinical context for the real-data evaluation."},{"cited_title":"E., Pollard, T","cited_arxiv_id":null,"evidence_quote":"Provides the MIMIC-III database from which the sepsis trajectories are extracted."}],"review_version":1}