Pith. sign in

REVIEW 3 major objections 5 minor 14 references

The orthocomplement of the tangent space for any Markov model is the direct sum of the orthocomplements of its single-constraint pieces, each given by a simple conditional-expectation formula.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 22:03 UTC pith:MQ2J4XBH

load-bearing objection Useful closed-form IF classes for non-DAG Markov models, but the load-bearing proof that T equals the intersection of the single-CI tangent spaces is broken as written. the 3 major comments →

arxiv 2607.23439 v1 pith:MQ2J4XBH submitted 2026-07-26 stat.ME cs.AI

A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models

classification stat.ME cs.AI MSC 62G0562H2262F12
keywords semiparametric efficiencyinfluence functionstangent spaceMarkov modelsconditional independencegraphical modelsADMGschain graphs
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Markov models are defined only by conditional independence restrictions. Efficient estimation of a finite-dimensional parameter inside such a model requires the class of all influence functions, which is one influence function plus the orthocomplement of the model’s tangent space. For DAG models that orthocomplement was already known; for undirected graphs, chain graphs and ordinary ADMG models it was not. The paper shows that any Markov model is the intersection of single-constraint models, so its tangent-space orthocomplement is simply the direct sum of the already-known single-constraint orthocomplements. Each piece is an explicit four-term conditional expectation, giving a closed-form description of every influence function. The same construction yields an iterative projection that produces a sequence of influence functions with non-increasing variance, improving any initial regular asymptotically linear estimator.

Core claim

For a Markov model defined by K conditional independences Xi ⊥ Yi | Zi, the orthocomplement of its tangent space equals the direct sum of the K single-constraint orthocomplements. Every element of that orthocomplement is therefore written in closed form as the sum over i of the operators Π(hi | T⊥i) = E[h|xi,yi,zi] - E[h|xi,zi] - E[h|yi,zi] + E[h|zi].

What carries the argument

Theorem 3: the identification T⊥ = ⊕ i T⊥i together with the explicit four-term projection formula for each single-constraint orthocomplement. This identity converts the geometric fact that the model is an intersection into a concrete recipe for every influence function.

Load-bearing premise

That the tangent space of the full intersection model is exactly the intersection of the individual closed tangent spaces, so that the orthocomplement is their direct sum.

What would settle it

Exhibit a concrete Markov model and a mean-zero square-integrable function that lies in every single-constraint tangent space yet fails to be a score of any parametric submodel that simultaneously satisfies all the model’s conditional independences.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Every influence function for any smooth target in a UG, CG or ordinary ADMG model is now obtained by adding an arbitrary sum of the four-term operators to one known influence function.
  • An iterative sequence of single-constraint projections produces influence functions of non-increasing variance, yielding more efficient RAL estimators from any initial one.
  • The same orthocomplement formula applies to non-graphical Markov models defined by arbitrary lists of conditional independences.
  • Once a projection onto the full orthocomplement is found, the efficient influence function itself becomes available for these models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The missing projection onto the full orthocomplement is now the only remaining obstacle to efficient influence functions; successive single-constraint projections may converge to it under additional regularity.
  • The same direct-sum geometry should extend immediately to models that also impose Verma constraints once the orthocomplement of a single Verma constraint is characterized.
  • Software that already computes single-constraint influence functions can be reused, without new algebraic factorization, to produce the full class for undirected and mixed-graph models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies semiparametric Markov models — statistical models defined solely by a list of conditional independence constraints Xi ⊥ Yi | Zi, i = 1,...,K — with emphasis on graphical models not equivalent to any DAG model (undirected graphs, chain graphs, ADMGs under the ordinary Markov property). The main result (Theorem 3) characterizes the orthocomplement of the model's tangent space as the direct sum of the orthocomplements of the K single-constraint models, each element written in closed form via the projection Π(h|T⊥i) = E[h|xi,yi,zi] − E[h|xi,zi] − E[h|yi,zi] + E[h|zi]. The argument views the model as an intersection of single-independence DAG models, identifies the tangent space with the intersection of the individual tangent spaces (T = ∩Ti), and invokes the Hilbert-space identity (∩Ti)⊥ = closure(ΣT⊥i) in the spirit of Van Der Laan & Robins (2003, Lemma 1.7). Applications include the class of influence functions for conditional-mean targets in the graphs of Figure 1 (§5.1) and an iterative efficiency-improvement scheme based on sequential projections onto the T⊥i (Theorem 4). The authors correctly note that the projection operator onto T⊥, and hence the efficient influence function, remains open.

Significance. If the main result holds, this is a useful and cleanly formulated contribution: the first closed-form characterization of the orthocomplement of the tangent space for ordinary Markov models associated with undirected graphs, chain graphs, and ADMGs, a class for which no explicit likelihood factorization is known and for which the DAG technique provably fails. The payoff is concrete: the full linear variety of influence functions for any pathwise-differentiable target (Eq. 28), an implementable variance-reduction iteration (Theorem 4) requiring only single-constraint projections in closed form, and instructive worked examples — notably the Bell scenario, where the one-step improvement of the saturated-model IF recovers the familiar AIPW form (Appendix D, Example 2). The paper is also commendably honest about what it does not deliver: the projection onto T⊥, the efficient influence function, and the implicitization of the likelihood remain open, and Example 1 in Appendix C explicitly flags that the natural candidate operator is not evidently a projection. These are genuine strengths. However, the characterization is parameter-free only if the identification T = ∩Ti is correct, and as

major comments (3)
  1. [Appendix C, proof of Theorem 3, second paragraph] The nontrivial inclusion T1∩...∩TK ⊆ T is not established. The proof takes f ∈ ∩Ti, sets pε = p0(1+εf), and asserts: 'Following the proof of Lemma 1, the i-th constraint must hold in pε(v).' This is not what Lemma 1's proof shows. Lemma 1 (Appendix B) verifies the constraint only for f drawn from each of the four factor subspaces TW|XYZ, TX|Z, TY|Z, TZ separately, where the perturbation factorizes through the conditional densities; for a general element of Ti (a sum of such pieces) the multiplicative perturbation breaks the constraint at O(ε²). Concrete counterexample: single constraint X⊥Y with Z,W empty, p0 uniform on {0,1}², f(x,y) = (−1)^x + (−1)^y. Then Π(f|T⊥) = E[f|x,y] − E[f|x] − E[f|y] + E[f] = 0, so f ∈ T (and indeed f is realizable by the factor-wise curve pε(x)pε(y) with pε(x) = (1+ε(−1)^x)/2). But the paper's curve gives pε(x,y) = (1+ε(−1)^x+ε(−1)^y)/4 while pε(x)pε(y) = (1+
  2. [§4.1 Lemma 2 and Appendix C, closedness of the sum of orthocomplements] The proofs assert that the sum T⊥1 + ... + T⊥K is closed and equals the direct sum ⊕i T⊥i ('Since both T⊥1 and T⊥2 are closed, T⊥1 ⊕ T⊥2 = overline{T⊥1 ⊕ T⊥2}'). A sum of two closed subspaces of a Hilbert space need not be closed; the correct general statement is (T1 ∩ T2)⊥ = closure(T⊥1 + T⊥2). The equality used is fine for finite state spaces (all subspaces finite-dimensional) but this assumption is never stated. Relatedly, the 'direct sum' terminology is misleading: the T⊥i need not be linearly independent. For instance, with the constraint list {X⊥Y, X⊥Y|Z} one has T⊥1 ⊆ T⊥2, and for the undirected square the two constraints both remove the same highest-order log-linear interaction, so the summands overlap. Equation (27) as a Minkowski sum is the correct object; the text should say 'sum', state the state-space assumption under which the sum is closed, and note the representation hi i
  3. [Global: regularity framework for all submodel constructions] The manuscript never specifies whether variables are discrete or general, whether p0 is required to have full support, or what class of scores is admitted. The construction pε = p0(1+εf) with 'δ small enough so that pε ≥ 0' (Lemma 8, Lemma 1, and Theorem 3 proofs) requires f bounded; for unbounded L²(p0) scores the standard truncation-plus-density argument is needed and should be written down once. Similarly, the reduction of pathwise differentiability (Eq. 4) to multiplicative submodels is flagged 'under regularity conditions' but the conditions are never stated, and conditional expectations E[·|z] are used at points where p0(z) may be small. These are routine in the semiparametric literature but must be collected into an explicit assumption set, especially since the paper's main claim is about general state spaces where, per Major Comment 2, the topological step also needs care.
minor comments (5)
  1. [§4.2 and §5.1, equation numbering] The main text repeatedly refers to 'Equation 81' (e.g., §4.2 after Theorem 3, and §5.1 'according to Equation 81'), which is the supplementary-material numbering of Eq. (27). Please unify the numbering or add cross-reference notes.
  2. [§5.2, Theorem 4] The iteration φm = φm−1 − Π(φm−1|T⊥im) with cyclic im is exactly the method of alternating (cyclic) projections onto the subspaces Ti. By von Neumann's theorem (K=2) and Halperin's extension (K≥2), φm converges to Π(φ0|T), i.e., to the efficient influence function, whenever T = ∩Ti and the relevant sums are closed. Citing this literature would answer, or at least sharply frame, the open question raised about the M→∞ limit, and would connect Theorem 4 to known convergence-rate results.
  3. [Appendix C, Example 1] The computation of ⟨(h−Γ(h)), Γ(g)⟩ is left inconclusive ('It is not evident that term2 = term1'). Using the Bell constraints at p0 (A⊥B,D and B⊥A,C), several cross terms factorize (e.g., E[E[h|A]E[g|B,D]] = E[E[h|A]·E[E[g|B,D]|A]]-type simplifications); it should be possible to either exhibit h, g with nonzero inner product (settling that Γ is not the projection) or prove equality. An inconclusive displayed computation should not be left in the supplement.
  4. [§5.1, notation for projection operators] In §5.1 the projection operators are subscripted Πa, Πb, Πc, then Πd for the bidirected square model P(e) of Figure 1e; the Bell scenario P(d) is treated in §4.1 instead. The lettering mismatch (Πd attached to model (e)) is confusing; consider subscripting by the model superscripts (a)–(e) or by constraint index.
  5. [Typos and references] Typos and small items: 'irrelevance,and' (§1); 'referred to asgraphical Markov models' (§1); 'important subsclass' (§2); 'BDs' in the local-Markov bullet list versus 'BGs' elsewhere (§2.1); 'syntatically' (§4.2); 'auxillary' (§5.1); 'The second one is is an iterative method' (§6); the equality sign in 'Its orthogonal complement is the direct sum' vs. sum (see Major Comment 2). Please also verify the numbering of the Van Der Laan & Robins (2003) lemma cited as Lemma 1.7.

Circularity Check

0 steps flagged

No circularity: the orthocomplement characterization is a direct Hilbert-space argument from single-constraint results and an intersection lemma, not a self-referential construction.

full rationale

The paper's central claim (Theorem 3) states that for a Markov model P = ∩ Pi defined by K CI constraints, T⊥ = ⊕ T⊥i with each T⊥i given by the closed-form projection of Lemma 1. This follows from the geometric identity that the orthocomplement of an intersection of closed subspaces is the sum of the orthocomplements (Van Der Laan & Robins Lemma 1.7, applied in the Appendix proof of Theorem 3), together with the already-known single-constraint orthocomplements. No parameter is fitted to data and then re-presented as a prediction; no uniqueness theorem is imported from the authors' prior work to forbid alternatives; the self-citations (Shpitser 2023, Bhattacharya et al. 2022, Tsiatis 2006) supply background factorization or DAG tangent-space facts that are independently established and are not load-bearing for the new intersection step. The derivation is therefore self-contained against its stated premises. (A separate correctness concern exists about whether the concrete submodel construction pε = p0(1+εf) stays inside the model for f ∈ ∩ Ti when K ≥ 2, but that is a gap in the proof of the inclusion, not a circular reduction of the claimed output to its inputs.)

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The result rests on standard L2 Hilbert-space geometry, the classical single-constraint orthocomplement, and the identification of the intersection model’s tangent space with the intersection of the component tangent spaces. No free parameters are fitted. No new physical or statistical entities are postulated.

axioms (5)
  • standard math L2(p0) mean-zero Hilbert space with inner product ⟨f,g⟩=E[fg]; tangent space is the L2-closure of scores of regular parametric submodels.
    Standard semi-parametric setup (Bickel et al., Tsiatis); used throughout §§2–4.
  • domain assumption For a single CI X⊥Y|Z the orthocomplement is {E[h|x,y,z]−E[h|x,z]−E[h|y,z]+E[h|z] : h∈H} (Lemma 1).
    Reproduced from Tsiatis Thm 4.5 / Rotnitzky–Smucler / Bhattacharya et al.; proved again in Appendix B.
  • domain assumption Tangent space of an intersection of models equals the intersection of the (closed) tangent spaces, hence T⊥ equals the direct sum of the component orthocomplements (Van Der Laan & Robins Lemma 1.7).
    Invoked as Lemma 2 and generalized in the proof of Theorem 3; the paper supplies an explicit submodel construction to justify T=∩Ti.
  • domain assumption Regularity conditions allowing pathwise differentiability and the von Mises expansion for the target functional ψ.
    Standard semi-parametric regularity (Tsiatis Ch. 3–4); assumed when moving from T⊥ to the IF class φ⊕T⊥.
  • ad hoc to paper Ordinary Markov ADMG models are defined solely by ordinary CI constraints (no Verma/generalized independence constraints).
    Explicitly restricted in §2.1; the characterization does not cover nested Markov models with Verma constraints.

pith-pipeline@v1.2.0-grok45-kimik3 · 35772 in / 2927 out tokens · 51930 ms · 2026-07-30T22:03:19.095807+00:00 · methodology

0 comments
read the original abstract

Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate finite-dimensional target parameters in such models efficiently, semi-parametric theory provides a principled framework for constructing regular and asymptotically linear estimators via influence functions (IFs). These estimators are asymptotically normal and root-$n$ consistent. Characterizing the class of all influence functions for a target parameter is crucial for statistically efficient inference in these models. For models that are Markov relative to directed acyclic graphs (DAGs), the orthogonal complement of the tangent space is known, implying that for any target the class of all influence functions can be derived once an influence function is obtained. On the other hand, for Markov models not equivalent to a DAG model -- such as ordinary Markov models associated with undirected graphs, chain graphs, or acyclic directed mixed graphs -- the orthogonal complement has not been characterized, impeding semi-parametric inference in these models. We derive closed form expressions for the orthogonal complement of the tangent space for general Markov models and illustrate our results by characterizing the class of influence functions for the conditional mean parameter in several graphical models.

Figures

Figures reproduced from arXiv: 2607.23439 by Ilya Shpitser, Trung Phung.

Figure 1
Figure 1. Figure 1: Six different graphs used to define six different [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: A graph representing a model with a Verma con [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 4 canonical work pages

  1. [4]

    One way to find the efficient influence function is projecting any other influence function φ onto the tangent space T

    The efficient influence function is the influence function φ(e) in the class of all influence functions φ⊕ T⊥ with smallest variance, i.e., E[(φ(e))2]≤E[φ ′2] for any influence function φ′ ∈φ⊕ T⊥ [Tsiatis, 2006]. One way to find the efficient influence function is projecting any other influence function φ onto the tangent space T . Let Π(· | T)be the proj...

  2. [5]

    doi: 10.1201/9780429463976

    ISBN 978-0-429-46397-6. doi: 10.1201/9780429463976. Whitney K. Newey. Semiparametric Efficiency Bounds. Journal of Applied Econometrics, 5(2):99–135,

  3. [7]

    Consider any disjoint subsets X,Y⊆V

    Lemma 5(Tsiatis [2006], Theorem 4.5).Let P be a semi-parametric model over variables V. Consider any disjoint subsets X,Y⊆V . For any parametric submodelpε, the score function s(x|y) = ∂ ∂ε logp ε(x|y)| ε=0 at p0 is an element of the subspace TX|Y :={E[h|x,y]−E[h|y] :∀h∈ H}.(38) Proof.LetZ=V\(X ˙∪Y). The following is a really useful identify for the score...

  4. [11]

    Let f be any function in T ⊥

    A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models (Supplementary Material) Trung Phung 1 Ilya Shpitser 1 1Computer Science Department, Johns Hopkins University, Baltimore, Maryland, USA A ADDITIONAL DETAILS ABOUT SEMIPARAMETRIC THEORY The set of all influence function and the orthogonal complement of the tangen...

  5. [14]

    First caseT A ⊆ T:Pick anyf=E[h|a]∈ T A. Then pε(c, b, a) = Z p0(d, c, b, a)(1 +εE[h|a])dd=p0(c, b, a)(1 +εE[h|a])(45) Therefore, by definition of conditional distribution pε(c|b, a) = pε(c, b, a) pε(b, a) =p 0(c|b, a) pε(c|a) = pε(c, a) pε(a) =p 0(c|a) pε(d|c, b, a) =pε(d, c, b, a) pε(c, b, a)=p 0(d|c, b, a) (46) This showsp ε(c|b, a) =p 0(c|b, a) =p 0(c...

  6. [1990]

    doi: 10.1002/jae.3950050202

    ISSN 1099-1255. doi: 10.1002/jae.3950050202. Elizabeth L. Ogburn, Ilya Shpitser, and Youjin Lee. Causal Inference, Social Networks and Chain Graphs.Journal of the Royal Statistical Society Series A: Statistics in Society, 183(4):1659–1676, October

  7. [2002]

    doi: 10.1111/1467-9868.00340

    ISSN 1467-9868. doi: 10.1111/1467-9868.00340. H. F. Lopes, E. Salazar, and D. Gamerman. Spatial dynamic factor analysis.Bayesian Analysis, 3(4):759–792,

  8. [2003]

    doi: 10.1007/978-0-387-21700-0

    ISBN 978-1-4419-3055-2 978-0-387-21700-0. doi: 10.1007/978-0-387-21700-0. A. W. van der Vaart.Asymptotic Statistics. Cambridge University Press, Cambridge,

  9. [2015]

    doi: 10.1007/978-3-319-16721-3

    ISBN 978-3-319-16720-6 978- 3-319-16721-3. doi: 10.1007/978-3-319-16721-3. Robin J. Evans. Margins of discrete bayesian networks. Annals of Statistics, 46:2623–2656,

  10. [2018]

    doi: 10.1111/ectj.12097

    ISSN 1368-4221. doi: 10.1111/ectj.12097. David A. Cox, John Little, and Donal O’Shea.Ideals, Vari- eties, and Algorithms: An Introduction to Computational Algebraic Geometry and Commutative Algebra. Under- graduate Texts in Mathematics. Springer International Publishing, Cham,

  11. [2019]

    doi: 10.3150/17-BEJ1005

    ISSN 1350-7265. doi: 10.3150/17-BEJ1005. M. Frydenberg. The chain graph Markov property.Scandi- navian Journal of Statistics,

  12. [2020]

    doi: 10.1111/rssa.12594

    ISSN 0964-1998. doi: 10.1111/rssa.12594. Judea Pearl.Probabilistic Reasoning in Intelligent Sys- tems: Networks of Plausible Inference. Morgan Kauf- mann, s.l.,

  13. [2021]

    doi: 10.1080/01621459.2020.1811098

    ISSN 0162-1459. doi: 10.1080/01621459.2020.1811098. Anastasios A. Tsiatis.Semiparametric Theory and Miss- ing Data. Springer Series in Statistics. Springer, New York, NY ,

  14. [2023]

    doi: 10.1214/22-AOS2253

    ISSN 0090-5364, 2168-8966. doi: 10.1214/22-AOS2253. Andrea Rotnitzky and Ezequiel Smucler. Efficient Adjust- ment Sets for Population Average Causal Treatment Ef- fect Estimation in Graphical Models.Journal of Machine Learning Research, 21(188):1–86,