{"id":"2c097ea1-dee0-4336-8d72-e0222104c8b5","arxiv_id":"2506.13025","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For a permutation missingness model, the authors derive an identifying expression and influence function for the mean of a partially missing outcome, enabling one-step efficient estimation.","lead":"This discussion of Nabi et al. (2022) proposes using single world intervention graphs for missing data and derives influence functions for estimating mean outcomes under a missing-not-at-random permutation model. It contributes efficient estimation tools that let analysts move beyond identification to estimation in MNAR problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The efficiency claim is the load-bearing new result, and it rests on an unproved Lemma 1 and unstated estimation conditions; the identifying assumptions are inherited and not the main risk.","rationale":"The reader's weakest_assumption points to the permutation-model independences: R1 independent of Y(1) given X(1), and R2 independent of (Y(1), X(1)) given Y and R1. These are indeed necessary for Proposition 1, but they are assumptions of the problem inherited from the original papers, and the discussion explicitly conditions on the permutation model. The genuinely new and contested part of the strongest_claim is the efficiency/one-step result. Since Lemma 1 is stated without proof, and no theorem supplies rates or conditions under which the remainder is negligible, the claim that the estimator is asymptotically efficient is not currently established. This is a support gap, not a refutation; the influence function passes basic mean-zero checks, the binary case is partially derived in the appendix, and the discussion paper format may tolerate a sketch. Therefore the appropriate disposition remains CONDITIONAL: the result may be correct, but it needs a completed derivation and stated conditions. I agree with the reader's overall verdict but locate the weakest point differently, hence partial agreement.","tokens_in":9124,"tokens_out":6827,"duration_ms":75641,"concrete_test":"Independently derive Lemma 1 by differentiating theta(P_epsilon) along a smooth one-dimensional submodel (for example, an exponential tilt of the observed-data law) and confirm that d theta(P_epsilon)/d epsilon at epsilon=0 equals the integral of phi(o; P) times the score, and that E_P[phi]=0. If the identity fails for a generic submodel, the influence function is wrong and the efficiency claim collapses. For the binary unknown-rho case, repeat the calculation treating rho as a functional of P, include its contribution to the influence function, and compare with the claimed 'extra asymptotically linear term'. A small symbolic or numerical check in a simple parametric permutation model would settle whether the displayed phi is correct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's genuinely new contribution is not identification, which is inherited from Robins (1997) and Nabi et al. (2022), but the claim that the functional theta admits the displayed influence function and that the displayed one-step estimator is asymptotically efficient. The most load-bearing step is Lemma 1, whose proof is explicitly omitted ('We omit details, but...'), and the subsequent assertion that plugging in a sqrt-n-consistent estimator of rho when rho is unknown contributes only an extra asymptotically linear term. No theorem states the regularity or rate conditions under which the remainder R_theta(P; Phat) is o_P(n^{-1/2}), nor verifies the empirical-process conditions needed for the one-step estimator to be asymptotically linear with influence function phi. The appendix derives only the binary, rho-known case, and that derivation uses an influence function for xi(x) stated for discrete X. Thus, as written, the central efficiency claim rests on an unverified calculation plus implicit high-level conditions. This is a support gap rather than a demonstrated error; the displayed phi does pass basic mean-zero checks, and the binary-case remainder is plausibly second-order. But if Lemma 1's phi fails the pathwise-derivative identity, the proposed estimator need not attain the claimed semiparametric efficiency bound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This discussion paper responds to Nabi et al. (2022) in three directions: (i) it surveys causal-inference tools (instrumental variables, shadow variables, negative controls) that could transfer to missing-data problems; (ii) it introduces the idea of m-SWIGs for simultaneous interventions on treatment and missingness, illustrated on a missing-exposure example; and (iii) it derives identification and estimation results for the permutation missingness model of Robins (1997), which is a primary example in Nabi et al. (2022). Specifically, Proposition 1 gives an identifying expression for ψ = E(Y(1)) in terms of observed-data functionals, Lemma 1 claims a von Mises expansion and influence function for θ = E(Y(1) | R1=0) in the general (non-binary) outcome case, and Corollary 2 gives a simplified influence function and a one-step estimator in the binary-outcome case with known odds ratio ρ. The paper argues that the proposed estimator is sqrt(n)-consistent and asymptotically efficient when ρ is known, and that plugging in a sqrt(n)-consistent estimator of ρ introduces only an extra asymptotically linear term.","tokens_in":9339,"tokens_out":2547,"duration_ms":31397,"significance":"If the efficiency results are correct, the paper provides the first practical semiparametric efficient estimator for a mean functional in the permutation missingness model, a non-standard MNAR model where estimation theory is otherwise underdeveloped. The m-SWIG discussion and the survey of transferable causal tools are useful conceptual contributions that will likely stimulate further work. The identification result in Proposition 1 is derived in detail and is a useful simplification of the full-law identification of Nabi et al. (2022). However, the central new technical contribution—the influence function and asymptotic-efficiency claim—is not fully established: Lemma 1's proof is omitted, and the extension to unknown ρ is only sketched. The efficiency claim therefore rests on unverified calculations rather than a complete theoretical argument. This is a support gap rather than a demonstrated error, and the paper is clearly written and well organized.","major_comments":[{"comment":"Lemma 1 is the load-bearing step for the paper's efficiency claim, but its proof is omitted: the manuscript says 'We omit details, but the result follows from calculations similar to those discussed for example in Section 4 of Kennedy (2022).' The influence function φ(O;P) must satisfy the pathwise derivative identity, and the remainder Rθ(P;P̄) must be o_P(n^{-1/2}) under suitable conditions, but neither is verified. As written, a reader cannot check the correctness of the displayed φ. Please provide a full proof (or at least a detailed derivation) of the von Mises expansion, including the mean-zero property of φ and a bound on the remainder that would make the one-step estimator asymptotically linear.","section":"Section 4, Lemma 1"},{"comment":"The influence function in Corollary 2 is derived under the assumption that ρ is known, but the proposed estimator then plugs in an estimated ρ̂, with the statement that 'the resulting estimator of θ will just have an extra asymptotically linear term.' No theorem is stated that gives conditions under which the one-step estimator with estimated ρ is asymptotically linear with a known influence function, nor is the extra term characterized. To support the efficiency claim, the authors should state a formal theorem that includes the regularity conditions (e.g., consistency and rate conditions on ξ̂ and ϖ̂, Donsker or empirical-process conditions, and the structure of the extra term from estimating ρ).","section":"Section 4, Corollary 2 and the one-step estimator"},{"comment":"The appendix derivation uses 'the influence function for ξ(x) when X is discrete.' The main results are presented for arbitrary X, including continuous covariates in the HIV example. For continuous X, the influence function for the conditional odds ξ(x) involves nonparametric estimation over a continuum, and the displayed remainder and one-step estimator require additional conditions (e.g., smoothness, rate conditions, or sample-splitting) that are not provided. This gap limits the generality of the efficiency claim to discrete X unless the authors supply the appropriate continuous-X theory.","section":"Appendix, proof of Corollary 2"}],"minor_comments":[{"comment":"In the displayed expression for ψ, the term θ is used before it is formally defined; please define θ = E(Y(1) | R1=0) immediately before the proposition to avoid confusion.","section":"Section 4, Proposition 1"},{"comment":"The interpretation of the posterior odds result is helpful, but the notation ζ(Y) is reused from Lemma 1 for a different quantity; consider a distinct symbol (e.g., ρ0(Y)) for the odds ratio to avoid ambiguity.","section":"Section 4, Corollary 1"}],"recommendation":"major_revision","confidential_remarks":"This is a discussion paper, so the bar for full proofs may be lower than for a regular research article. Nevertheless, the efficiency result is a central part of the contribution, and the omission of Lemma 1's proof and the sketchy treatment of unknown ρ are substantive gaps. I would recommend major revision, with the expectation that the authors provide a complete proof or clearly mark the result as a conjecture, and that they add a precise theorem for the plug-in estimator with unknown ρ. The paper otherwise makes a valuable contribution to the discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Levis and Kennedy's discussion piece makes one genuinely new contribution: an influence function for the functional θ = E(β(X)/α(X) | R1=0,R2=1) in Robins's permutation model, plus a one-step estimator and a clean binary-case simplification. The identification formula for ψ in Proposition 1 is a repackaging of Nabi et al.'s full-law identification, and the authors say so. The new mathematics starts at Lemma 1, and that is the part that needs scrutiny.\n\nThe paper does several things well. The early sections lay out, in a readable way, how shadow variables, instrumental variables, and SWIGs transfer from causal inference to missing data. The binary case is genuinely elegant: θ becomes a posterior odds calculation with a clear interpretation in the HIV example. The proposed estimator is concrete, and the appendix works through the binary remainder carefully enough to make the second-order terms plausible.\n\nThe soft spots are real but not fatal. Lemma 1 is the load-bearing result, and its proof is omitted, deferred to calculations similar to Kennedy (2022). That is honest but unsatisfying: there is no theorem stating the conditions under which the remainder is o_P(n^{-1/2}), and no verification of the empirical-process or rate conditions for the one-step estimator. Corollary 2 derives the influence function for known ρ; the unknown-ρ case gets a single sentence claiming that a √n-consistent plug-in just adds an asymptotically linear term. The derivation of the influence function for ξ(x) is stated for discrete X, and extending to continuous X is not discussed. These are support gaps, not demonstrated errors. The displayed φ passes basic mean-zero checks, and the binary remainder derivation in the appendix gives me some confidence the expansion is right.\n\nThe citation pattern is fine: the paper builds on Robins (1997) and Nabi et al. (2022), and the new claims are not circular. No fitting is done and no parameters are tuned to data, so there is no circularity concern.\n\nWho should read this: people working on efficient estimation in MNAR models, especially those interested in permutation-model identification, and readers of the original Nabi et al. paper. It deserves a serious referee. A referee can ask for the missing proof of Lemma 1, the regularity conditions for the remainder, and a proper treatment of the unknown-ρ and continuous-X cases. That is the right way to turn a plausible expansion into a reliable result.\n\nI would send this to peer review.","headline":"A useful discussion with a real new influence function for the permutation MNAR model, but the efficiency claim is under-supported as written.","tokens_in":9865,"tokens_out":2869,"would_cite":true,"duration_ms":27530,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D10","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under the permutation missingness model, the mean of a partially missing variable can be estimated at the parametric rate with local efficiency, and binary outcomes reduce to a posterior odds calculation.","keywords":["missing not at random","permutation missingness model","influence function","semiparametric efficiency","one-step estimation","counterfactual independence","single world intervention graphs"],"falsifier":"Generate data from a model that satisfies both independence restrictions, compute the proposed one-step estimator with correctly specified nuisance functions on many samples, and check that the standardized estimator is approximately standard normal; a systematic deviation would disprove the influence function.","tokens_in":8907,"feed_emoji":"📊","tokens_out":7779,"duration_ms":75082,"temperature":0.7,"pith_summary":"This discussion argues that the causal and counterfactual framework of the paper under discussion can be pushed in two directions: borrowing additional causal identification tools for missing-not-at-random problems, and constructing efficient estimators for target functionals rather than the full data law. Its central concrete result concerns the permutation missingness model: the mean $\\psi = E(Y(1))$ of the partially missing variable is identified by a ratio of conditional expectations, and the piece $\\theta = E(Y(1) \\mid R_1 = 0)$ satisfies a von Mises expansion with an explicit influence function. That expansion yields a one-step estimator that is $\\sqrt{n}$-consistent and locally efficient, so a researcher who trusts the model's independence restrictions does not need to estimate the entire observed-data distribution. When $Y$ is binary, the estimand simplifies to a posterior odds calculation, and the influence function becomes a simple product of odds and density ratios.","feed_headline":"Missing data mean gets an efficient estimator under permutation MNAR","feed_subtitle":"A closed-form influence function gives root-n-consistent, locally efficient estimation; binary outcomes reduce to a posterior odds…","key_machinery":"The load-bearing object is the permutation missingness model's pair of nonparametric independence restrictions: $R_1 \\perp\\!\\!\\perp Y(1) \\mid X(1)$ and $R_2 \\perp\\!\\!\\perp (Y(1), X(1)) \\mid Y, R_1$. These restrictions identify the full data law and, via Bayes' rule, imply the ratio formula for $\\theta$. The analytical engine is the von Mises expansion of the functional $\\theta = E(\\beta(X)/\\alpha(X) \\mid R_1 = 0, R_2 = 1)$, whose influence function is derived in Lemma 1; the remainder term is explicitly displayed, making the second-order rate transparent. In the binary case the key simplification is the odds identity $\\xi(X) = \\lambda(X) \\times \\text{odds}(Y = 1 \\mid R_1 = R_2 = 1)$, where $\\lambda$ is a density ratio, so the mean functional becomes the posterior probability of $Y = 1$ from prior odds times likelihood ratio, standardized to the $R_1 = 0, R_2 = 1$ group.","core_discovery":"Under the permutation missingness model, the discussion establishes that the target mean $\\psi = E(Y(1))$ is identified from observed data $(X, R, Y)$ by $\\psi = P(R_1 = 1)E(Y \\mid R_1 = 1) + P(R_1 = 0)\\theta$, where $\\theta$ is a ratio of conditional expectations involving $\\zeta(Y) = P(R_2 = 1 \\mid R_1 = 1, Y)$. Proposition 1 states this identification, and Lemma 1 shows that $\\theta$ has the von Mises expansion $\\theta(P) - \\theta(P) = \\int \\varphi(o; P) \\, d(P - P)(o) + R_\\theta(P; P)$ with the displayed influence function $\\varphi(O; P)$; this gives the local asymptotic minimax lower bound. For binary $Y$, $\\theta = E\\{\\xi(X)/(\\rho + \\xi(X)) \\mid R_1 = 0, R_2 = 1\\}$, where $\\xi$ is a conditional odds and $\\rho$ an odds ratio, interpreted as a posterior odds combining prior information from the $R_1 = 1$ group with a likelihood ratio from the doubly observed group. Corollary 2 provides the corresponding influence function in the binary case, and the resulting one-step estimator is displayed.","pith_inferences":["The same ratio-of-expectations structure suggests that double/debiased machine learning with flexible nuisance estimators could be applied directly; the discussion only sketches the one-step estimator.","The posterior odds interpretation implies the estimand is monotone in the prior odds and in the density ratio $\\lambda(x)$, which could be used to construct sensitivity bounds under partial violations of the independence assumptions.","The explicit remainder in the von Mises expansion shows products of second-order nuisance errors, the structure required for rate double robustness; the discussion does not develop this point.","A natural extension is to models with more than two missingness indicators, where the same permutation-style independence restrictions would yield recursive ratio formulas; the paper does not state this."],"forward_implications":["Under the permutation model, estimating $\\psi$ does not require estimating the joint distribution of $(X(1), Y(1))$; the closed-form ratio in Proposition 1 suffices.","The influence function in Lemma 1 gives the efficiency bound for $\\theta$, so the one-step estimator attains local asymptotic minimax optimality.","With binary $Y$, the estimand has a direct posterior odds reading, making the missing-not-at-random adjustment transparent to practitioners.","Replacing the unknown odds ratio $\\rho$ by its plug-in estimate preserves $\\sqrt{n}$-consistency and asymptotic normality, adding only an extra asymptotically linear term.","SWIGs (m-SWIGs) can deliver counterfactual independence conditions such as $A(1) \\perp\\!\\!\\perp (Y^{a(1)}, R^{a(1)}) \\mid X$ that are not directly visible from m-DAG d-separation."],"supporting_citations":[{"why":"Introduces the permutation missingness model whose independence restrictions identify the full law and underpin Proposition 1.","marker":"Robins (1997)"},{"why":"The discussed paper; supplies the graphical identification of the full data law that the discussion extends to functionals.","marker":"Nabi et al. (2022)"},{"why":"Provides the von Mises expansion and influence function techniques used in Lemma 1 and Corollary 2.","marker":"Kennedy (2022)"},{"why":"Introduces SWIGs used in Section 3 to derive counterfactual independences involving missingness.","marker":"Richardson and Robins (2013)"},{"why":"Supplies the Bayes-rule identity for missing exposure data used in the m-SWIG example.","marker":"Kennedy (2020)"}],"fun_headline_variants":["Permutation MNAR mean gets efficient estimator via causal tools","Causal perspective yields closed-form influence function for MNAR","Efficient estimation for missing-not-at-random mean identified causally","Discussion shows root-n consistent estimator for permutation MNAR","Causal and counterfactual views unlock MNAR identification and estimation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The permutation model assumes that $R_1$ is unrelated to the unobserved outcome given the first covariates, and that $R_2$ is unrelated to both unobserved variables once $Y$ and $R_1$ are known; if either is false, the identifying formula and the efficient estimator are invalid.","fun_headline_variants_meta":{"raw":{"variants":["Permutation MNAR mean gets efficient estimator via causal tools","Causal perspective yields closed-form influence function for MNAR","Efficient estimation for missing-not-at-random mean identified causally","Discussion shows root-n consistent estimator for permutation MNAR","Causal and counterfactual views unlock MNAR identification and estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1363,"prompt_tokens":996,"completion_tokens":367,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":284}},"tokens_in":612,"tokens_out":367,"duration_ms":4695,"temperature":1.0,"reasoning_tokens":284,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:35:59.981207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data from a model that satisfies both independence restrictions, compute the proposed one-step estimator with correctly specified nuisance functions on many samples, and check that the standardized estimator is approximately standard normal; a systematic deviation would disprove the influence function.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the permutation missingness model whose independence restrictions identify the full law and underpin Proposition 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces SWIGs used in Section 3 to derive counterfactual independences involving missingness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Bayes-rule identity for missing exposure data used in the m-SWIG example."}],"review_version":1}